Online Learning of Rested and Restless Bandits
From MaRDI portal
Abstract: In this paper we study the online learning problem involving rested and restless multiarmed bandits with multiple plays. The system consists of a single player/user and a set of K finite-state discrete-time Markov chains (arms) with unknown state spaces and statistics. At each time step the player can play M arms. The objective of the user is to decide for each step which M of the K arms to play over a sequence of trials so as to maximize its long term reward. The restless multiarmed bandit is particularly relevant to the application of opportunistic spectrum access (OSA), where a (secondary) user has access to a set of K channels, each of time-varying condition as a result of random fading and/or certain primary users' activities.
Cited in
(6)- An online algorithm for the risk-aware restless bandit
- A Survey of Preference-Based Online Learning with Bandit Algorithms
- Normal bandits of unknown means and variances
- Game of thrones: fully distributed learning for multiplayer bandits
- Interior-Point Methods for Full-Information and Bandit Online Learning
- On the sensitivity of restless bandits solutions to uncertainty in the models of the arms
This page was built for publication: Online Learning of Rested and Restless Bandits
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2989865)