TD-regularized actor-critic methods
From MaRDI portal
Publication:2320580
Abstract: Actor-critic methods can achieve incredible performance on difficult reinforcement learning problems, but they are also prone to instability. This is partly due to the interaction between the actor and critic during learning, e.g., an inaccurate step taken by one of them might adversely affect the other and destabilize the learning. To avoid such issues, we propose to regularize the learning objective of the actor by penalizing the temporal difference (TD) error of the critic. This improves stability by avoiding large steps in the actor update whenever the critic is highly inaccurate. The resulting method, which we call the TD-regularized actor-critic method, is a simple plug-and-play approach to improve stability and overall performance of the actor-critic methods. Evaluations on standard benchmarks confirm this.
Recommendations
Cites work
- \(\text{Q}(\lambda)\) with off-policy corrections
- scientific article; zbMATH DE number 6982305 (Why is no real title available?)
- scientific article; zbMATH DE number 2107836 (Why is no real title available?)
- scientific article; zbMATH DE number 5060482 (Why is no real title available?)
- OnActor-Critic Algorithms
- Reinforcement learning. An introduction
- Simple statistical gradient-following algorithms for connectionist reinforcement learning
- Variance reduction techniques for gradient estimates in reinforcement learning
Cited in
(7)- Improve generated adversarial imitation learning with reward variance regularization
- Hyperbolically Discounted Temporal Difference Learning
- td-reg
- A Small Gain Analysis of Single Timescale Actor Critic
- Optimistic reinforcement learning by forward Kullback-Leibler divergence optimization
- On the sample complexity of actor-critic method for reinforcement learning with function approximation
- Actor prioritized experience replay
This page was built for publication: TD-regularized actor-critic methods
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2320580)