Powered by OpenAIRE graph
Found an issue? Give us feedback
addClaim

Restless bandits: indexability, computation of whittle index and learning

Authors: Akbarzadeh, Nima;

Restless bandits: indexability, computation of whittle index and learning

Abstract

Les bandits agités sont une classe de problèmes d'allocation séquentielle de ressources concernés par l'allocation d'une ou plusieurs ressources entre plusieurs processus alternatifs où l'évolution des processus est markovienne et dépend des ressources qui leur sont allouées. En 1988, Whittle a développé une heuristique d'indice pour les problèmes de bandits agités qui est devenue une approche de solution populaire en raison de sa simplicité et de ses solides performances empiriques. L'heuristique de l'indice de Whittle est applicable si le modèle satisfait une condition technique connue sous le nom d'indexabilité.Dans cette thèse, nous nous concentrons sur trois configurations générales : les modèles entièrement observables, partiellement observables et les modèles d'apprentissage.Pour la configuration entièrement observable, nous présentons deux conditions générales suffisantes pour l'indexabilité et identifions des raffinements plus simples pour vérifier ces conditions. Ensuite, nous revisitons un algorithme proposé précédemment appelé algorithme glouton adaptatif qui est connu pour calculer l'indice de Whittle pour une sous-classe de bandits agités. Nous montrons qu'une généralisation de l'algorithme glouton adaptatif calcule l'indice de Whittle pour tous les bandits agités indexables. Nous présentons une implémentation efficace de cet algorithme qui peut calculer l'indice de Whittle d'un bras avec $K$ états dans $\OO(K^3)$ calculs. Enfin, nous présentons une étude numérique détaillée qui confirme les bonnes performances de l'heuristique de l'indice de Whittle.En ce qui concerne la configuration partiellement observable, nous considérons un redémarrage des bandits avec deux modèles d'observation : premièrement, l'état de chaque bandit n'est pas du tout observable, et deuxièmement, l'état de chaque bandit n'est observable que s'il est choisi. Pour les deux modèles, nous montrons que le système est indexable. Pour le premier modèle, nous dérivons une expression de forme fermée pour l'indice de Whittle. Pour le deuxième modèle, nous proposons un algorithme efficace pour calculer l'indice de Whittle en exploitant les propriétés qualitatives de la politique optimale. Nous présentons des expériences numériques détaillées qui indiquent que la politique d'indice de Whittle surpasse la politique myope et peut être proche de l'optimum dans différentes configurations.Pour la configuration d'apprentissage, nous supposons que la véritable dynamique du système est inconnue et présentons un algorithme d'apprentissage par renforcement (RL) d'échantillonnage de Thompson qui a un regret qui évolue de manière polynomiale avec le nombre de bras. Cela contraste avec la plupart des algorithmes RL où le regret cumulé évolue de manière exponentielle avec le nombre de bras. Nous présentons deux caractérisations du regret de l'algorithme proposé par rapport à la politique d'indice de Whittle. En fonction du modèle de récompense et des hypothèses techniques, nous montrons que pour un bandit agité avec $n$ bras et au plus $m$ activations à chaque fois, le regret évolue soit en $\tilde{\mathcal{O}}(n^2 \sqrt{T})$, $\tilde{\mathcal{O}}(mn\sqrt{T})$, $\tilde{\mathcal{O}}(n^{1.5}\sqrt{T})$, ou $\tilde{\mathcal{O}}(\max\{m\sqrt{n}, n\}\sqrt{T})$. Nous présentons des exemples numériques pour illustrer les principales caractéristiques de l'algorithme

Restless bandits are a class of sequential resource allocation problems concerned with allocating one or more resources among several alternative processes where the evolution of the processes is Markovian and depends on the resources allocated to them. In 1988, Whittle developed an index heuristic for restless bandit problems which has emerged as a popular solution approach due to its simplicity and strong empirical performance. The Whittle index heuristic is applicable if the model satisfies a technical condition known as indexability. In this thesis, we focus on three general setups: fully-observable, partially-observable, and learning models. For the fully-observable setup, we present two general sufficient conditions for indexability and identify simpler to verify refinements of these conditions. Afterwards, we revisit a previously proposed algorithm called adaptive greedy algorithm which is known to compute the Whittle index for a subclass of restless bandits. We show that a generalization of the adaptive greedy algorithm computes the Whittle index for all indexable restless bandits. We present an efficient implementation of this algorithm which can compute the Whittle index of an arm with $K$ states in $\OO(K^3)$ computations. Finally, we present a detailed numerical study which affirms the strong performance of the Whittle index heuristic.Regarding the partially-observable setup, we consider a restart bandits with two observational models: first, the state of each bandit is not observable at all, and second, the state of each bandit is observable only if it is chosen. For both models, we show that the system is indexable. For the first model, we derive a closed-form expression for the Whittle index. For the second model, we propose an efficient algorithm to compute the Whittle index by exploiting the qualitative properties of the optimal policy. We present detailed numerical experiments which indicates that the Whittle index policy outperforms myopic policy and can be close to optimal in different setups.For the learning setup, we assume the true system dynamics are unknown and present a Thompson sampling reinforcement learning (RL) algorithm which has a regret which scales polynomially with the number of arms. This is in contrast to most RL algorithms where the cumulative regret scales exponentially with the number of arms. We present two characterizations of the regret of the proposed algorithm with respect to the Whittle index policy. Depending on the reward model and technical assumptions, we show that for a restless bandit with $n$ arms and at most $m$ activations at each time, the regret scales either as $\tilde{\mathcal{O}}(n^2 \sqrt{T})$, $\tilde{\mathcal{O}}(mn\sqrt{T})$, $\tilde{\mathcal{O}}(n^{1.5}\sqrt{T})$, or $\tilde{\mathcal{O}}(\max\{m\sqrt{n}, n\}\sqrt{T})$. We present numerical examples to illustrate the salient features of the algorithm

Mahajan, Aditya (Supervisor)

Country
Canada
Related Organizations
Keywords

Electrical and Computer Engineering

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Related to Research communities
Upload OA version
Are you the author of this publication? Upload your Open Access version to Zenodo!
It’s fast and easy, just two clicks!