Optimal adaptive policies for sequential allocation problems
From MaRDI portal
Recommendations
- Asymptotically efficient adaptive allocation rules
- Optimal sequential allocation with imperfect feedback information
- Optimal Adaptive Policies for Markov Decision Processes
- scientific article; zbMATH DE number 4003938
- scientific article; zbMATH DE number 4045510
- Adaptive Sequential Stochastic Optimization
- Dynamic allocation policies for the finite horizon one armed bandit problem
Cited in
(40)- On bidding for a fixed number of items in a sequence of auctions
- Robust control of the multi-armed bandit problem
- The multi-armed bandit problem: an efficient nonparametric solution
- Adaptive policies for perimeter surveillance problems
- Asymptotically optimal algorithms for budgeted multiple play bandits
- Reading policies for joins: an asymptotic analysis
- Consistency of sequential Bayesian sampling policies
- Self-Optimizing and Pareto-Optimal Policies in General Environments based on Bayes-Mixtures
- Response-adaptive designs for clinical trials: simultaneous learning from multiple patients
- Optimal sequential sampling from two populations.
- Kullback-Leibler upper confidence bounds for optimal sequential allocation
- Adaptive aggregation for reinforcement learning in average reward Markov decision processes
- Robustness of stochastic bandit policies
- An asymptotically optimal policy for finite support models in the multiarmed bandit problem
- scientific article; zbMATH DE number 1538064 (Why is no real title available?)
- Optimal and asymptotically optimal decision rules for sequential screening and resource allocation
- scientific article; zbMATH DE number 6982311 (Why is no real title available?)
- Normal bandits of unknown means and variances
- Learning the distribution with largest mean: two bandit frameworks
- Structured Policies for a Sequential Design Problem with General Distributions
- On large deviations properties of sequential allocation problems
- Exploration-exploitation policies with almost sure, arbitrarily slow growing asymptotic regret
- Infinite Arms Bandit: Optimality via Confidence Bounds
- Sequential Bayes-optimal policies for multiple comparisons with a known standard
- Explore first, exploit next: the true shape of regret in bandit problems
- Adaptive policies for sequential sampling under incomplete information and a cost constraint
- Multi-armed bandits under general depreciation and commitment
- Asymptotically optimal multi-armed bandit policies under a cost constraint
- Irreversible adaptive allocation rules
- Tracking the mean of a piecewise stationary sequence
- Pair-matching: link prediction with adaptive queries
- Tracking the mean of a piecewise stationary sequence
- An optimal selection for ensembles of influential projects
- Uniform mean estimation for monotonic processes
- CRIMED: lower and upper bounds on regret for bandits with unbounded stochastic corruption
- On best-arm identification with a fixed budget in non-parametric multi-armed bandits
- Bandit algorithms based on Thompson sampling for bounded reward distributions
- Optimal -correct best-arm selection for heavy-tailed distributions
- A non-parametric solution to the multi-armed bandit problem with covariates
- A perpetual search for talents across overlapping generations: a learning process
This page was built for publication: Optimal adaptive policies for sequential allocation problems
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q1922542)