The sample complexity of exploration in the multi-armed bandit problem
From MaRDI portal
Recommendations
- Lower bounds on the sample complexity of exploration in the multi-armed bandit problem.
- scientific article; zbMATH DE number 2089367
- Sample mean based index policies by O(log n) regret for the multi-armed bandit problem
- Finite-time analysis of the multiarmed bandit problem
- The Nonstochastic Multiarmed Bandit Problem
Cited in
(41)- A PAC algorithm in relative precision for bandit problem with costly sampling
- Trading utility and uncertainty: applying the value of information to resolve the exploration-exploitation dilemma in reinforcement learning
- Sequential estimation of quantiles with applications to A/B testing and best-arm identification
- Good arm identification via bandit feedback
- Pure exploration in finitely-armed and continuous-armed bandits
- On the complexity of best-arm identification in multi-armed bandit models
- Approximation algorithms for stochastic combinatorial optimization problems
- scientific article; zbMATH DE number 2089367 (Why is no real title available?)
- Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
- Bayesian Incentive-Compatible Bandit Exploration
- Online Regret Bounds for Markov Decision Processes with Deterministic Transitions
- Pure exploration in multi-armed bandits problems
- Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
- The \(K\)-armed dueling bandits problem
- Finite-time lower bounds for the two-armed bandit problem
- Learning the distribution with largest mean: two bandit frameworks
- Pure exploration in infinitely-armed bandit models with fixed-confidence
- Near-optimal PAC bounds for discounted MDPs
- Sample mean based index policies by O(log n) regret for the multi-armed bandit problem
- Nonasymptotic sequential tests for overlapping hypotheses applied to near-optimal arm identification in bandit models
- Preference-based online learning with dueling bandits: a survey
- On the bias, risk, and consistency of sample means in multi-armed bandits
- Amplification and Derandomization without Slowdown
- Tractable sampling strategies for ordinal optimization
- Simple Bayesian algorithms for best-arm identification
- Best arm identification for contaminated bandits
- Explore first, exploit next: the true shape of regret in bandit problems
- Lower bounds on the sample complexity of exploration in the multi-armed bandit problem.
- Lower bounds and selectivity of weak-consistent policies in stochastic multi-armed bandit problem
- An instance-based algorithm for deciding the bias of a coin
- UCB revisited: improved regret bounds for the stochastic multi-armed bandit problem
- Certified multifidelity zeroth-order optimization
- Semiparametric inference based on adaptively collected data
- Optimizing Sharpe ratio: risk-adjusted decision-making in multi-armed bandits
- On the problem of best arm retention
- On the problem of best arm retention
- Top-k combinatorial bandits with full-bandit feedback
- Episodic reinforcement learning in finite MDPs: minimax lower bounds revisited
- Multi-armed bandits with episode context
- A perpetual search for talents across overlapping generations: a learning process
- Online regret bounds for Markov decision processes with deterministic transitions
This page was built for publication: The sample complexity of exploration in the multi-armed bandit problem
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3093197)