Sample mean based index policies by O(log n) regret for the multi-armed bandit problem
From MaRDI portal
Publication:4862097
Recommendations
- On the bias, risk, and consistency of sample means in multi-armed bandits
- An index-based deterministic convergent optimal algorithm for constrained multi-armed bandit problems
- On an index policy for restless bandits
- The sample complexity of exploration in the multi-armed bandit problem
- Lower bounds on the sample complexity of exploration in the multi-armed bandit problem.
- Index-based policies for discounted multi-armed bandits on parallel machines.
Cited in
(56)- Exploration-exploitation tradeoff using variance estimates in multi-armed bandits
- On Bayesian index policies for sequential resource allocation
- An improved upper bound on the expected regret of UCB-type policies for a matching-selection bandit problem
- Efficient crowdsourcing of unknown experts using bounded multi-armed bandits
- An online algorithm for the risk-aware restless bandit
- A revised approach for risk-averse multi-armed bandits under CVaR criterion
- Gittins' theorem under uncertainty
- Multi-agent reinforcement learning: a selective overview of theories and algorithms
- The multi-armed bandit problem: an efficient nonparametric solution
- How fragile are information cascades?
- Multi-armed bandits based on a variant of simulated annealing
- Geiringer theorems: from population genetics to computational intelligence, memory evolutive systems and Hebbian learning
- Non-asymptotic analysis of a new bandit algorithm for semi-bounded rewards
- Approximate indexability and bandit problems with concave rewards and delayed feedback
- The sample complexity of exploration in the multi-armed bandit problem
- Deviations of stochastic bandit regret
- Linearly parameterized bandits
- Tuning Bandit Algorithms in Stochastic Environments
- Wisdom of crowds versus groupthink: learning in groups and in isolation
- Kullback-Leibler upper confidence bounds for optimal sequential allocation
- Exploration and exploitation of scratch games
- Robustness of stochastic bandit policies
- An asymptotically optimal policy for finite support models in the multiarmed bandit problem
- Some memoryless bandit policies
- Finite-time lower bounds for the two-armed bandit problem
- scientific article; zbMATH DE number 6982311 (Why is no real title available?)
- Normal bandits of unknown means and variances
- Learning the distribution with largest mean: two bandit frameworks
- Finite-time analysis for the knowledge-gradient policy
- Boundary crossing for general exponential families
- Nonasymptotic Analysis of Monte Carlo Tree Search
- Optimistic Gittins Indices
- Exploration-exploitation policies with almost sure, arbitrarily slow growing asymptotic regret
- Infinite Arms Bandit: Optimality via Confidence Bounds
- Continuous Assortment Optimization with Logit Choice Probabilities and Incomplete Information
- Explore first, exploit next: the true shape of regret in bandit problems
- Derivative-free optimization methods
- Lower bounds on the sample complexity of exploration in the multi-armed bandit problem.
- Lower bounds and selectivity of weak-consistent policies in stochastic multi-armed bandit problem
- Functional Sequential Treatment Allocation
- Dealing with expert bias in collective decision-making
- Convergence rate analysis for optimal computing budget allocation algorithms
- Empirical Gittins index strategies with -explorations for multi-armed bandit problems
- A confirmation of a conjecture on Feldman’s two-armed bandit problem
- UCB revisited: improved regret bounds for the stochastic multi-armed bandit problem
- Factorial Designs for Online Experiments
- Adaptive maximization of social welfare
- Domain independent heuristics for online stochastic contingent planning
- The multi-armed bandit problem under the mean-variance setting
- Simple fixes that accommodate switching costs in multi-armed bandits
- Output-weighted sampling for multi-armed bandits with extreme payoffs
- \textsc{Athanor}: local search over abstract constraint specifications
- Decentralized cooperative reinforcement learning with hierarchical information structure
- Boundary crossing probabilities for general exponential families
- Contextual combinatorial conservative bandits
- A non-parametric solution to the multi-armed bandit problem with covariates
This page was built for publication: Sample mean based index policies by O(log n) regret for the multi-armed bandit problem
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4862097)