Rollout sampling approximate policy iteration
From MaRDI portal
Recommendations
Cites work
- scientific article; zbMATH DE number 3148886 (Why is no real title available?)
- 10.1162/1532443041827907
- Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
- Approximate policy iteration with a policy language bias: solving relational Markov decision processes
- Finite-time analysis of the multiarmed bandit problem
- Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
Cited in
(9)- An approximate policy iteration viewpoint of actor-critic algorithms
- An Incremental Fast Policy Search Using a Single Sample Path
- Nonparametric approximation generalized policy iteration reinforcement learning algorithm based on states clustering
- Machine Learning: ECML 2004
- Efficient exploration through active learning for value function approximation in reinforcement learning
- A survey of preference-based reinforcement learning methods
- Analysis of classification-based policy iteration algorithms
- Preference-based reinforcement learning: a formal framework and a policy iteration algorithm
- Dynamic parcel pick-up routing problem with prioritized customers and constrained capacity via lower-bound-based rollout approach
This page was built for publication: Rollout sampling approximate policy iteration
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2036256)