The Nonstochastic Multiarmed Bandit Problem
From MaRDI portal
Recommendations
- The multi-armed bandit problem: an efficient nonparametric solution
- The Irrevocable Multiarmed Bandit Problem
- Multi-armed bandit problem revisited
- The Multi-Armed Bandit With Stochastic Plays
- The Multi-Armed Bandit Problem: Decomposition and Computation
- scientific article; zbMATH DE number 4084786
- Optimal exploration-exploitation in a multi-armed bandit problem with non-stationary rewards
- The Continuum-Armed Bandit Problem
- A Structured Multiarmed Bandit Problem and the Greedy Policy
Cites work
- Algorithms – ESA 2005
- Efficient crowdsourcing of unknown experts using multi-armed bandits
- Eliminating spammers and ranking annotators for crowdsourced labeling tasks
- scientific article; zbMATH DE number 3567782 (Why is no real title available?)
- Optimal aggregation of classifiers in statistical learning.
- The Nonstochastic Multiarmed Bandit Problem
Cited in
(only showing first 100 items - show all)- Exploration-exploitation tradeoff using variance estimates in multi-armed bandits
- Perspectives on multiagent learning
- Multi-agent learning for engineers
- Further contributions to the two-armed bandit problem
- Strategy under the unknown stochastic environment: The nonparametric lop-pass problem
- Improving multi-armed bandit algorithms in online pricing settings
- Two queues with non-stochastic arrivals
- Randomized prediction of individual sequences
- Selective harvesting over networks
- Online learning in online auctions
- Online multiple kernel classification
- Extracting certainty from uncertainty: regret bounded by variation in costs
- Regret bounds for sleeping experts and bandits
- Bayesian adversarial multi-node bandit for optimal smart grid protection against cyber attacks
- Gorthaur-EXP3: bandit-based selection from a portfolio of recommendation algorithms balancing the accuracy-diversity dilemma
- Multi-armed bandit with sub-exponential rewards
- Gittins' theorem under uncertainty
- Dismemberment and design for controlling the replication variance of regret for the multi-armed bandit
- Stochastic continuum-armed bandits with additive models: minimax regrets and adaptive algorithm
- Multi-agent reinforcement learning: a selective overview of theories and algorithms
- Trading utility and uncertainty: applying the value of information to resolve the exploration-exploitation dilemma in reinforcement learning
- Robust control of the multi-armed bandit problem
- MedleySolver: online SMT algorithm selection
- Adaptive large neighborhood search for mixed integer programming
- The multi-armed bandit problem: an efficient nonparametric solution
- Dynamic pricing with finite price sets: a non-parametric approach
- Filtered Poisson process bandit on a continuum
- Tune and mix: learning to rank using ensembles of calibrated multi-class classifiers
- BoostingTree: parallel selection of weak learners in boosting, with application to ranking
- Mistake bounds on the noise-free multi-armed bandit game
- New bounds on the price of bandit feedback for mistake-bounded online multiclass learning
- Analysis of Hannan consistent selection for Monte Carlo tree search in simultaneous move games
- A bad arm existence checking problem: how to utilize asymmetric problem structure?
- On the stability of an adaptive learning dynamics in traffic games
- Exponential weight approachability, applications to calibration and regret minimization
- Improved second-order bounds for prediction with expert advice
- Online calibrated forecasts: memory efficiency versus universality for learning in games
- Global Nash convergence of Foster and Young's regret testing
- Pure exploration in finitely-armed and continuous-armed bandits
- Online linear optimization and adaptive routing
- Reinforcement learning and evolutionary algorithms for non-stationary multi-armed bandit problems
- Multi-armed bandits based on a variant of simulated annealing
- Mechanisms with learning for stochastic multi-armed bandit problems
- Doubly robust policy evaluation and optimization
- On two continuum armed bandit problems in high dimensions
- Value functions for depth-limited solving in zero-sum imperfect-information games
- Regret minimization in online Bayesian persuasion: handling adversarial receiver's types under full and partial feedback models
- Multi-channel transmission scheduling with hopping scheme under uncertain channel states
- Truthful mechanisms with implicit payment computation
- Noise free multi-armed bandit game
- Batched bandit problems
- On the Prior Sensitivity of Thompson Sampling
- Online learning in Markov decision processes with continuous actions
- Algorithms for computing strategies in two-player simultaneous move games
- Close the gaps: a learning-while-doing algorithm for single-product revenue management problems
- Achieving Unbounded Resolution inFinitePlayer Goore Games Using Stochastic Automata, and Its Applications
- Playing in stochastic environment: from multi-armed bandits to two-player games
- Discount targeting in online social networks using backpressure-based learning
- Learning where to attend with deep architectures for image tracking
- Chasing Ghosts: Competing with Stateful Policies
- Agent-based Modeling and Simulation of Competitive Wholesale Electricity Markets
- Sequential decision making with vector outcomes
- Computational Randomness from Generalized Hardcore Sets
- The sample complexity of exploration in the multi-armed bandit problem
- On upper-confidence bound policies for switching bandit problems
- The Irrevocable Multiarmed Bandit Problem
- scientific article; zbMATH DE number 7038557 (Why is no real title available?)
- On learning algorithms for Nash equilibria
- No regret learning in oligopolies: Cournot vs. Bertrand
- Bandit online optimization over the permutahedron
- Generalized mirror descents in congestion games
- On Solving Finite State Multi-Armed Bandit Problem by Linear Programming
- Reinforcement with fading memories
- Bayesian Incentive-Compatible Bandit Exploration
- Incentivizing exploration with heterogeneous value of money
- Following the Perturbed Leader to Gamble at Multi-armed Bandits
- A Simple Distribution-Free Approach to the Max k-Armed Bandit Problem
- Online Regret Bounds for Markov Decision Processes with Deterministic Transitions
- Workspace-based connectivity oracle: an adaptive sampling strategy for PRM planning
- Gaussian process modelling of dependencies in multi-armed bandit problems
- Pure exploration in multi-armed bandits problems
- Exploration and exploitation of scratch games
- Algorithm portfolio selection as a bandit problem with unbounded losses
- An asymptotically optimal policy for finite support models in the multiarmed bandit problem
- Combinatorial bandits
- Learning with stochastic inputs and adversarial outputs
- The \(K\)-armed dueling bandits problem
- Beyond the hazard rate: more perturbation algorithms for adversarial multi-armed bandits
- Bandits with knapsacks
- Nonstochastic Multi-Armed Bandits with Graph-Structured Feedback
- Bandit regret scaling with the effective loss range
- Delay and cooperation in nonstochastic bandits
- Thompson sampling guided stochastic searching on the line for deceptive environments with applications to root-finding problems
- Better algorithms for benign bandits
- A minimax and asymptotically optimal algorithm for stochastic bandits
- Regret bounds for restless Markov bandits
- Learning Theory
- Sequential Shortest Path Interdiction with Incomplete Information
- The Nonstochastic Multiarmed Bandit Problem
- scientific article; zbMATH DE number 1907146 (Why is no real title available?)
This page was built for publication: The Nonstochastic Multiarmed Bandit Problem
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4785631)