Bandit algorithms
From MaRDI portal
Stopping times; optimal stopping problems; gambling theory (60G40) Optimal stopping in statistics (62L15) Research exposition (monographs, survey articles) pertaining to computer science (68-02) Learning and adaptive systems in artificial intelligence (68T05) Pattern recognition, speech recognition (68T10) Markov and semi-Markov decision processes (90C40) Probabilistic games; gambling (91A60)
Recommendations
Cited in
(only showing first 100 items - show all)- A penalized bandit algorithm
- Regularizing double machine learning in partially linear endogenous models
- Randomized allocation with nonparametric estimation for contextual multi-armed bandits with delayed rewards
- Multi-objective multi-armed bandit with lexicographically ordered and satisficing objectives
- Uncertainty calibration for probabilistic projection methods
- Multi-armed bandit with sub-exponential rewards
- Two-armed bandit problem and batch version of the mirror descent algorithm
- A PAC algorithm in relative precision for bandit problem with costly sampling
- Stochastic continuum-armed bandits with additive models: minimax regrets and adaptive algorithm
- Fundamental design principles for reinforcement learning algorithms
- A Markov decision process for response-adaptive randomization in clinical trials
- Whittle index based Q-learning for restless bandits with average reward
- Decision making under uncertainty and reinforcement learning. Theory and algorithms
- Deciding when to quit the gambler's ruin game with unknown probabilities
- Ballooning multi-armed bandits
- On concentration inequalities for vector-valued Lipschitz functions
- Budget-limited distribution learning in multifidelity problems
- Bayesian adaptive randomization with compound utility functions
- Customization of J. Bather's UCB strategy for a Gaussian multiarmed bandit
- Batched bandit problems
- Negatively Correlated Bandits
- Sequential learning and decision-making in wireless resource management
- Matching while learning
- Preference-based online learning with dueling bandits: a survey
- On multi-armed bandit designs for dose-finding trials
- scientific article; zbMATH DE number 7370627 (Why is no real title available?)
- On the bias, risk, and consistency of sample means in multi-armed bandits
- A bandit-learning approach to multifidelity approximation
- scientific article; zbMATH DE number 7596797 (Why is no real title available?)
- scientific article; zbMATH DE number 7626761 (Why is no real title available?)
- Bayesian exploration: incentivizing exploration in Bayesian games
- Bayesian brains and the Rényi divergence
- A Single-Index Model With a Surface-Link for Optimizing Individualized Dose Rules
- Multiplayer Bandits Without Observing Collision Information
- Fictitious play in zero-sum stochastic games
- Always Valid Inference: Continuous Monitoring of A/B Tests
- Locks, Bombs and Testing: The Case of Independent Locks
- Individual fairness in hindsight
- Are we forgetting about compositional optimisers in Bayesian optimisation?
- Achieving fairness in the stochastic multi-armed bandit problem
- Multi-Armed Bandits: Theory and Applications to Online Learning in Networks
- Introduction to multi-armed bandits
- scientific article; zbMATH DE number 6253908 (Why is no real title available?)
- Learning in repeated auctions
- Bypassing the Monster: A Faster and Simpler Optimal Algorithm for Contextual Bandits Under Realizability
- Robust sequential design for piecewise-stationary multi-armed bandit problem in the presence of outliers
- Functional Sequential Treatment Allocation
- Greedy Algorithm Almost Dominates in Smoothed Contextual Bandits
- Daisee: Adaptive importance sampling by balancing exploration and exploitation
- Online learning for scheduling MIP heuristics
- Multi-armed bandit-based hyper-heuristics for combinatorial optimization problems
- Risk filtering and risk-averse control of Markovian systems subject to model uncertainty
- Multi-armed bandits with censored consumption of resources
- Dealing with expert bias in collective decision-making
- Nonparametric learning for impulse control problems -- exploration vs. exploitation
- A Theory of Bounded Inductive Rationality
- A unified stochastic approximation framework for learning in games
- A probabilistic reduced basis method for parameter-dependent problems
- Learning Stationary Nash Equilibrium Policies in n-Player Stochastic Games with Independent Chains
- UCB strategies and optimization of batch processing in a one-armed bandit problem
- Nearly Dimension-Independent Sparse Linear Bandit over Small Action Spaces via Best Subset Selection
- Constrained regret minimization for multi-criterion multi-armed bandits
- AI-driven liquidity provision in OTC financial markets
- Asymptotic optimality for decentralised bandits
- Temporal logic explanations for dynamic decision systems using anchors and Monte Carlo tree search
- Safe multi-agent reinforcement learning for multi-robot control
- Treatment recommendation with distributional targets
- Exponential asymptotic optimality of Whittle index policy
- Response-adaptive randomization in clinical trials: from myths to practical considerations
- Empirical Gittins index strategies with -explorations for multi-armed bandit problems
- Efficient and generalizable tuning strategies for stochastic gradient MCMC
- Robust and efficient algorithms for conversational contextual bandit
- Relaxing the i.i.d. assumption: adaptively minimax optimal regret via root-entropic regularization
- Settling the sample complexity of model-based offline reinforcement learning
- Analyzing bandit-based adaptive operator selection mechanisms
- Exploiting action impact regularity and exogenous state variables for offline reinforcement learning
- Optimistic MLE: a generic model-based algorithm for partially observable sequential decision making
- Finding the optimal exploration-exploitation trade-off online through Bayesian risk estimation and minimization
- Thompson sampling for networked control over unknown channels
- Adaptive Algorithm for Multi-Armed Bandit Problem with High-Dimensional Covariates
- Invariant description of control in a Gaussian one-armed bandit problem
- Multinomial Thompson sampling for rating scales and prior considerations for calibrating uncertainty
- Optimal analysis for bandit learning in matching markets with serial dictatorship
- A modified EXP3 in adversarial bandits with multi-user delayed feedback
- Optimization of two-alternative batch processing with parameter estimation based on data inside batches
- Tracking the mean of a piecewise stationary sequence
- Online learning in budget-constrained dynamic Colonel Blotto games
- Surveillance for endemic infectious disease outbreaks: adaptive sampling using profile likelihood estimation
- Algorithms for decision making
- An exact bandit model for the risk-volatility tradeoff
- Certified multifidelity zeroth-order optimization
- Doubly Robust Interval Estimation for Optimal Policy Evaluation in Online Learning
- Deep spatial Q-learning for infectious disease control
- Thompson sampling-based recursive block elimination for dynamic assignment under limited budget in pure-exploration
- Risk preferences of learning algorithms
- Integrating multi-armed bandit with local search for MaxSAT
- Online learning in sequential Bayesian persuasion: handling unknown priors
- An approximate control variates approach to multifidelity distribution estimation
- Tracking the mean of a piecewise stationary sequence
- Bandit learning in matching markets with relative feedback
This page was built for publication: Bandit algorithms
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5109247)