Continuous-time allocation indices and their discrete-time approximation
From MaRDI portal
allocationoptimal expected rewardsoptimal policy in multi-armed bandit problemsstrong Markov processes
Stopping times; optimal stopping problems; gambling theory (60G40) Applications of Markov chains and discrete-time Markov processes on general state spaces (social mobility, learning theory, industrial processes, etc.) (60J20) Dynamic programming (90C39) Markov and semi-Markov decision processes (90C40)
Recommendations
Cited in
(12)- Continuous multi-armed bandits and multiparameter processes
- Multi-armed bandits in discrete and continuous time
- Discrete multiarmed bandits and multiparameter processes
- Dynamic allocation problems in continuous time
- Synchronization and optimality for multi-armed bandit problems in continuous time
- On Gittins' index theorem in continuous time
- Sensitivity of the gittins index in the contiuous time two-armed bandit problem
- A general theory of multiarmed bandit processes with constrained arm switches
- Open bandit processes with uncountable states and time-backward effects
- Empirical Gittins index strategies with -explorations for multi-armed bandit problems
- Lévy bandits under Poissonian decision times
- Gittins indices in the dynamic allocation problem for diffusion processes
This page was built for publication: Continuous-time allocation indices and their discrete-time approximation
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3745006)