Feel-Good Thompson Sampling for Contextual Bandits and Reinforcement Learning
From MaRDI portal
(Redirected from Publication:5089723)
Recommendations
Cites work
- 10.1162/153244303321897663
- A decision-theoretic generalization of on-line learning and an application to boosting
- An information-theoretic analysis of Thompson sampling
- Bypassing the Monster: A Faster and Simpler Optimal Algorithm for Contextual Bandits Under Realizability
- Competitive On-line Statistics
- Deep exploration via randomized value functions
- Information-theoretic determination of minimax rates of convergence
- Learning to optimize via posterior sampling
- The Nonstochastic Multiarmed Bandit Problem
Cited in
(6)- On the Prior Sensitivity of Thompson Sampling
- Thompson Sampling for Bayesian Bandits with Resets
- A Tutorial on Thompson Sampling
- Bayesian design principles for frequentist sequential learning
- Unified algorithms for RL with decision-estimation coefficients: PAC, reward-free, preference-based learning and beyond
- Adaptive network approach to exploration-exploitation trade-off in reinforcement learning
This page was built for publication: Feel-Good Thompson Sampling for Contextual Bandits and Reinforcement Learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5089723)