Interior-Point Methods for Full-Information and Bandit Online Learning
From MaRDI portal
Recommendations
- Online bandit convex optimisation with stochastic constraints via two-point feedback
- Stochastic online optimization. Single-point and multi-point non-linear multi-armed bandits. Convex and strongly-convex case
- Tsallis-INF: an optimal algorithm for stochastic and adversarial bandits
- Stochastic convex optimization with bandit feedback
- A bandit-learning approach to multifidelity approximation
- Online convex optimization in the bandit setting: gradient descent without a gradient
- Regret lower bound and optimal algorithm for high-dimensional contextual linear bandit
- Regret and Convergence Bounds for a Class of Continuum-Armed Bandit Problems
- An Optimal Algorithm for Bandit and Zero-Order Convex Optimization with Two-Point Feedback
- Online Learning of Rested and Restless Bandits
Cited in
(8)- On two continuum armed bandit problems in high dimensions
- No regret learning in oligopolies: Cournot vs. Bertrand
- Generalized mirror descents in congestion games
- Efficient sampling from time-varying log-concave distributions
- A generalized online mirror descent with applications to classification and regression
- Privacy-preserving distributed projected one-point bandit online optimization over directed graphs
- Online self-concordant and relatively smooth minimization, with applications to online portfolio selection and learning quantum states
- Learning equilibria in matching markets with bandit feedback
This page was built for publication: Interior-Point Methods for Full-Information and Bandit Online Learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5271795)