The Nonstochastic Multiarmed Bandit Problem
From MaRDI portal
Recommendations
- The multi-armed bandit problem: an efficient nonparametric solution
- The Irrevocable Multiarmed Bandit Problem
- Multi-armed bandit problem revisited
- The Multi-Armed Bandit With Stochastic Plays
- The Multi-Armed Bandit Problem: Decomposition and Computation
- scientific article; zbMATH DE number 4084786
- Optimal exploration-exploitation in a multi-armed bandit problem with non-stationary rewards
- The Continuum-Armed Bandit Problem
- A Structured Multiarmed Bandit Problem and the Greedy Policy
Cites work
- scientific article; zbMATH DE number 3567782 (Why is no real title available?)
- Algorithms – ESA 2005
- Efficient crowdsourcing of unknown experts using multi-armed bandits
- Eliminating spammers and ranking annotators for crowdsourced labeling tasks
- Optimal aggregation of classifiers in statistical learning.
- The Nonstochastic Multiarmed Bandit Problem
Cited in
(only showing first 100 items - show all)- Stochastic bandits robust to adversarial corruptions
- The Continuum-Armed Bandit Problem
- Doubly robust policy evaluation and optimization
- Pure exploration in multi-armed bandits problems
- Regret bounds for Narendra-Shapiro bandit algorithms
- Learning Theory
- Achieving Unbounded Resolution inFinitePlayer Goore Games Using Stochastic Automata, and Its Applications
- Online learning in online auctions
- Keyword-level Bayesian online bid optimization for sponsored search advertising
- On the Prior Sensitivity of Thompson Sampling
- A minimax and asymptotically optimal algorithm for stochastic bandits
- X-armed bandits
- On two continuum armed bandit problems in high dimensions
- Extracting certainty from uncertainty: regret bounded by variation in costs
- Competitive collaborative learning
- Replicator dynamics: old and new
- Crowdsourcing label quality: a theoretical analysis
- scientific article; zbMATH DE number 7370545 (Why is no real title available?)
- Feel-Good Thompson Sampling for Contextual Bandits and Reinforcement Learning
- scientific article; zbMATH DE number 7561642 (Why is no real title available?)
- scientific article; zbMATH DE number 7306905 (Why is no real title available?)
- Algorithm portfolio selection as a bandit problem with unbounded losses
- No regret learning in oligopolies: Cournot vs. Bertrand
- Per-round knapsack-constrained linear submodular bandits
- A natural adaptive process for collective decision-making
- Thompson sampling for networked control over unknown channels
- Multi-channel transmission scheduling with hopping scheme under uncertain channel states
- Pure exploration in finitely-armed and continuous-armed bandits
- Online calibrated forecasts: memory efficiency versus universality for learning in games
- Nonstationary bandits with habituation and recovery dynamics
- Selective harvesting over networks
- A reinforcement learning approach for resolving inconsistencies in qualitative constraint networks
- Constant regret for sequence prediction with limited advice
- Bypassing the Monster: A Faster and Simpler Optimal Algorithm for Contextual Bandits Under Realizability
- Finite-time analysis of the multiarmed bandit problem
- Simple fixes that accommodate switching costs in multi-armed bandits
- The rate of convergence of Bregman proximal methods: local geometry versus regularity versus sharpness
- Playing in stochastic environment: from multi-armed bandits to two-player games
- On learning algorithms for Nash equilibria
- Batched bandit problems
- Multi-Armed Bandits: Theory and Applications to Online Learning in Networks
- Algorithms for computing strategies in two-player simultaneous move games
- Learning in combinatorial optimization: what and how to explore
- The multi-armed bandit problem: an efficient nonparametric solution
- Regret minimization in online Bayesian persuasion: handling adversarial receiver's types under full and partial feedback models
- Bandits with knapsacks
- Bandit-based task assignment for heterogeneous crowdsourcing
- On the discrete-time origins of the replicator dynamics: from convergence to instability and chaos
- Exponential weight algorithm in continuous time
- Sequential interdiction with incomplete information and learning
- Computational Randomness from Generalized Hardcore Sets
- scientific article; zbMATH DE number 7380836 (Why is no real title available?)
- Combinatorial bandits
- Close the gaps: a learning-while-doing algorithm for single-product revenue management problems
- Improved algorithms for bandit with graph feedback via regret decomposition
- Beyond the hazard rate: more perturbation algorithms for adversarial multi-armed bandits
- Sequential decision making with vector outcomes
- Nonparametric bandit methods
- Reinforcement learning and evolutionary algorithms for non-stationary multi-armed bandit problems
- Small-Loss Bounds for Online Learning with Partial Information
- AI-driven liquidity provision in OTC financial markets
- Learning dynamic algorithm portfolios
- Algorithms for adversarial bandit problems with multiple plays
- UCB revisited: improved regret bounds for the stochastic multi-armed bandit problem
- Dynamic pricing with finite price sets: a non-parametric approach
- Filtered Poisson process bandit on a continuum
- Further contributions to the two-armed bandit problem
- Setting Reserve Prices in Second-Price Auctions with Unobserved Bids
- BoostingTree: parallel selection of weak learners in boosting, with application to ranking
- Online mixed discrete and continuous optimization: algorithms, regret analysis and applications
- Unified algorithms for RL with decision-estimation coefficients: PAC, reward-free, preference-based learning and beyond
- Two-armed restless bandits with imperfect information: stochastic control and indexability
- Bayesian adversarial multi-node bandit for optimal smart grid protection against cyber attacks
- Fare inspection patrolling under in-station selective inspection policy
- Global Nash convergence of Foster and Young's regret testing
- Learning with stochastic inputs and adversarial outputs
- The \(K\)-armed dueling bandits problem
- Adversarial contextual bandits go kernelized
- Slowly changing adversarial bandit algorithms are efficient for discounted MDPs
- Importance-weighted offline learning done right
- CRIMED: lower and upper bounds on regret for bandits with unbounded stochastic corruption
- Workspace-based connectivity oracle: an adaptive sampling strategy for PRM planning
- Multi-agent learning for engineers
- Two queues with non-stochastic arrivals
- A Simple Distribution-Free Approach to the Max k-Armed Bandit Problem
- Integrated online learning and adaptive control in queueing systems with uncertain payoffs
- Best-of-both-worlds algorithms for partial monitoring
- Adversarial online multi-task reinforcement learning
- Improved high-probability regret for adversarial bandits with time-varying feedback graphs
- Follow-the-perturbed-leader achieves best-of-both-worlds for bandit problems
- Online learning for traffic navigation in congested networks
- Online learning with off-policy feedback
- A unified algorithm for stochastic path problems
- Robust control of the multi-armed bandit problem
- Robust sequential design for piecewise-stationary multi-armed bandit problem in the presence of outliers
- Improving multi-armed bandit algorithms in online pricing settings
- Top-k combinatorial bandits with full-bandit feedback
- Thompson sampling for adversarial bit prediction
- First-order Bayesian regret analysis of Thompson sampling
- An asymptotically optimal policy for finite support models in the multiarmed bandit problem
This page was built for publication: The Nonstochastic Multiarmed Bandit Problem
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4785631)