Analysis of classification-based policy iteration algorithms
From MaRDI portal
Recommendations
- Rollout sampling approximate policy iteration
- Finite-sample analysis of least-squares policy iteration
- Approximate policy iteration: a survey and some new methods
- Regularized policy iteration with nonparametric function spaces
- Performance Loss Bounds for Approximate Value Iteration with State Aggregation
Cited in
(8)- Rollout sampling approximate policy iteration
- Performance guarantees for policy learning
- Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithm
- Safe policy iteration: a monotonically improving approximate policy iteration approach
- On the theory of policy gradient methods: optimality, approximation, and distribution shift
- A priori estimates for deep residual network in continuous-time reinforcement learning
- Deep approximate policy iteration
- Deep controlled learning for inventory control
This page was built for publication: Analysis of classification-based policy iteration algorithms
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2810787)