Adaptive learning of uncontrolled restless bandits with logarithmic regret

descriptionPublicationkeyboard_double_arrow_right Article , Conference object 01 Sep 2011Publisher:IEEEJournal:2011 49th Annual Allerton Conference on Communication, Control, and Computing (Allerton)

Authors: Cem Tekin; Mingyan Liu;

doi: 10.1109/allerton.2011.6120273

Adaptive learning of uncontrolled restless bandits with logarithmic regret

- Summary
- Metrics

Abstract

In this paper we consider the problem of learning the optimal policy for the uncontrolled restless bandit problem. In this problem only the state of the selected arm can be observed, the state transitions are independent of control and the transition law is unknown. We propose a learning algorithm which gives logarithmic regret uniformly over time with respect to the optimal finite horizon policy with known transition law under some assumptions on the transition probabilities of the arms and the structure of the optimal stationary policy for the infinite horizon average reward problem.

Related Organizations

University of Michigan–Flint
United States
University of Michigan–Ann Arbor
United States
University of Michigan
United States

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	12
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%