Tsallis-INF: an optimal algorithm for stochastic and adversarial bandits
From MaRDI portal
Publication:4998901
Recommendations
Cites work
- A generalized online mirror descent with applications to classification and regression
- Asymptotically efficient adaptive allocation rules
- Elements of Information Theory
- Finite-time analysis of the multiarmed bandit problem
- Kullback-Leibler upper confidence bounds for optimal sequential allocation
- Multi-player bandits revisited
- Online learning and online convex optimization
- Perturbation techniques in online learning and optimization
- Possible generalization of Boltzmann-Gibbs statistics.
- Prediction, Learning, and Games
- Regret analysis of stochastic and nonstochastic multi-armed bandit problems
- Regret bounds and minimax policies under partial monitoring
- Some aspects of the sequential design of experiments
- Stochastic bandits robust to adversarial corruptions
- The Nonstochastic Multiarmed Bandit Problem
- Thompson sampling: an asymptotically optimal finite-time analysis
Cited in
(11)- Improved regret for zeroth-order adversarial bandit convex optimisation
- Interior-Point Methods for Full-Information and Bandit Online Learning
- scientific article; zbMATH DE number 7064063 (Why is no real title available?)
- Online team formation under different synergies
- Relaxing the i.i.d. assumption: adaptively minimax optimal regret via root-entropic regularization
- Implicitly normalized forecaster with clipping for linear and non-linear heavy-tailed multi-armed bandits
- Adaptive maximization of social welfare
- Best-of-both-worlds algorithms for partial monitoring
- Follow-the-perturbed-leader achieves best-of-both-worlds for bandit problems
- Scale-free adversarial multi armed bandits
- Information-directed sampling for bandits: a primer
This page was built for publication: Tsallis-INF: an optimal algorithm for stochastic and adversarial bandits
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4998901)