On-line policy gradient estimation with multi-step sampling
From MaRDI portal
Recommendations
- scientific article; zbMATH DE number 1961138
- scientific article; zbMATH DE number 1753153
- A basic formula for performance gradient estimation of semi-Markov decision processes
- Approximate gradient methods in policy-space optimization of Markov reward processes
- Performance gradient estimation for the very large finite Markov chains
Cites work
- A basic formula for online policy gradient algorithms
- scientific article; zbMATH DE number 3532286 (Why is no real title available?)
- scientific article; zbMATH DE number 700091 (Why is no real title available?)
- scientific article; zbMATH DE number 1753152 (Why is no real title available?)
- scientific article; zbMATH DE number 1753153 (Why is no real title available?)
- Perturbation realization, potentials, and sensitivity analysis of Markov processes
- Simulation-based optimization of Markov reward processes
- Variance reduction techniques for gradient estimates in reinforcement learning
Cited in
(3)
This page was built for publication: On-line policy gradient estimation with multi-step sampling
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5962027)