An approximate policy iteration viewpoint of actor-critic algorithms
From MaRDI portal
Cites work
- 10.1162/1532443041827907
- \(\text{Q}(\lambda)\) with off-policy corrections
- \({\mathcal Q}\)-learning
- A review of stochastic algorithms with continuous value function approximation and some new approximate policy iteration algorithms for multidimensional continuous applications
- An analysis of temporal-difference learning with function approximation
- An emphatic approach to the problem of off-policy temporal-difference learning
- Approximate policy iteration: a survey and some new methods
- Asynchronous stochastic approximation and Q-learning
- Convergence of entropy-regularized natural policy gradient with linear function approximation
- Fast global convergence of natural policy gradient methods with entropy regularization
- Finite-sample analysis of least-squares policy iteration
- Finite-sample analysis of nonlinear stochastic approximation with applications in reinforcement learning
- Finite-Sample Analysis of Two-Time-Scale Natural Actor–Critic Algorithm
- scientific article; zbMATH DE number 1321699 (Why is no real title available?)
- scientific article; zbMATH DE number 1753152 (Why is no real title available?)
- Markov chains and mixing times. With a chapter on ``Coupling from the past by James G. Propp and David B. Wilson.
- Natural actor-critic algorithms
- On linear and super-linear convergence of natural policy gradient algorithm
- On the sample complexity of actor-critic method for reinforcement learning with function approximation
- On the theory of policy gradient methods: optimality, approximation, and distribution shift
- Policy mirror descent for reinforcement learning: linear convergence, new sampling complexity, and generalized problem classes
- Projected equation methods for approximate solution of large linear systems
- Reinforcement learning. An introduction
- Rollout sampling approximate policy iteration
- The actor-critic algorithm as multi-time-scale stochastic approximation.
This page was built for publication: An approximate policy iteration viewpoint of actor-critic algorithms
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6947638)