A Survey of Preference-Based Online Learning with Bandit Algorithms
From MaRDI portal
Recommendations
- Preference-based online learning with dueling bandits: a survey
- A tractable online learning algorithm for the multinomial logit contextual bandit
- A survey on online learning methods: Thompson sampling and others
- A survey of algorithms and analysis for adaptive online learning
- A survey of preference-based reinforcement learning methods
- Multi-Armed Bandits: Theory and Applications to Online Learning in Networks
- An Online Learning Approach to a Multi-player N-armed Functional Bandit
- Online Learning of Rested and Restless Bandits
- An Efficient Algorithm for Learning with Semi-bandit Feedback
- An online algorithm for the risk-aware restless bandit
Cited in
(8)- Interactive Thompson sampling for multi-objective multi-armed bandits
- Top-\(\kappa\) selection with pairwise comparisons
- Interactive preference elicitation under noisy preference models: an efficient non-Bayesian approach
- scientific article; zbMATH DE number 7370524 (Why is no real title available?)
- Efficient and optimal algorithms for contextual dueling bandits under realizability
- A survey of preference-based reinforcement learning methods
- Preference-based reinforcement learning: a formal framework and a policy iteration algorithm
- Query complexity of tournament solutions
This page was built for publication: A Survey of Preference-Based Online Learning with Bandit Algorithms
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2938721)