Bandit algorithms based on Thompson sampling for bounded reward distributions
From MaRDI portal
Cites work
- Asymptotically efficient adaptive allocation rules
- Finite-time analysis of the multiarmed bandit problem
- Kullback-Leibler upper confidence bounds for optimal sequential allocation
- Non-asymptotic analysis of a new bandit algorithm for semi-bounded rewards
- Optimal adaptive policies for sequential allocation problems
- Thompson sampling: an asymptotically optimal finite-time analysis
This page was built for publication: Bandit algorithms based on Thompson sampling for bounded reward distributions
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q7025113)