The actor-critic algorithm as multi-time-scale stochastic approximation.
From MaRDI portal
Cites work
- scientific article; zbMATH DE number 48727 (Why is no real title available?)
- scientific article; zbMATH DE number 51132 (Why is no real title available?)
- scientific article; zbMATH DE number 3538599 (Why is no real title available?)
- scientific article; zbMATH DE number 700091 (Why is no real title available?)
- scientific article; zbMATH DE number 941166 (Why is no real title available?)
- scientific article; zbMATH DE number 3232230 (Why is no real title available?)
- A tutorial survey of reinforcement learning
- Asynchronous stochastic approximation and Q-learning
- Chaotic relaxation
- Do stochastic algorithms avoid traps?
- Estimation and control in discounted stochastic dynamic programming
- Feature-based methods for large scale dynamic programming
- Generalized polynomial approximations in Markovian decision processes
- New method of stochastic approximation type
- Nonconvergence to unstable points in urn models and stochastic approximations
- Stochastic approximation methods for constrained and unconstrained systems
- Stochastic approximation with two time scales
- \({\mathcal Q}\)-learning
Cited in
(4)- An approximate policy iteration viewpoint of actor-critic algorithms
- Deep reinforcement learning for infinite horizon mean field problems in continuous spaces
- Stochastic approximation and reinforcement learning: the interface and a little beyond
- Reinforcement learning based algorithms for average cost Markov decision processes
This page was built for publication: The actor-critic algorithm as multi-time-scale stochastic approximation.
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5955801)