Satisficing in Time-Sensitive Bandit Learning
From MaRDI portal
Recommendations
- Thompson sampling: an asymptotically optimal finite-time analysis
- Satisficing: a 'pretty good' heuristic
- Multi-objective multi-armed bandit with lexicographically ordered and satisficing objectives
- A learning algorithm for the finite-time two-armed bandit problem
- An information-theoretic analysis of Thompson sampling
Cites work
- X-armed bandits
- A Tutorial on Thompson Sampling
- An information-theoretic analysis of Thompson sampling
- Asymptotically efficient adaptive allocation rules
- Bandit problems with infinitely many arms
- Choosing a good toolkit. I: Prior-free heuristics
- Choosing a good toolkit. II: Bayes-rule based heuristics
- scientific article; zbMATH DE number 5485582 (Why is no real title available?)
- Kullback-Leibler upper confidence bounds for optimal sequential allocation
- Learning to optimize via posterior sampling
- Linearly parameterized bandits
- On the complexity of best-arm identification in multi-armed bandit models
- Regret analysis of stochastic and nonstochastic multi-armed bandit problems
- The knowledge gradient algorithm for a general class of online learning problems
Cited in
(3)
This page was built for publication: Satisficing in Time-Sensitive Bandit Learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5870357)