New Versions of Gradient Temporal-Difference Learning
From MaRDI portal
Abstract: Sutton, Szepesv'{a}ri and Maei introduced the first gradient temporal-difference (GTD) learning algorithms compatible with both linear function approximation and off-policy training. The goal of this paper is (a) to propose some variants of GTDs with extensive comparative analysis and (b) to establish new theoretical analysis frameworks for the GTDs. These variants are based on convex-concave saddle-point interpretations of GTDs, which effectively unify all the GTDs into a single framework, and provide simple stability analysis based on recent results on primal-dual gradient dynamics. Finally, numerical comparative analysis is given to evaluate these approaches.
Recommendations
- Differential Temporal Difference Learning
- Technical update: Least-squares temporal difference learning
- Factored temporal difference learning in the New Ties environment
- Hyperbolically Discounted Temporal Difference Learning
- Generalized TD learning
- Practical issues in temporal difference learning
- Linear least-squares algorithms for temporal difference learning
This page was built for publication: New Versions of Gradient Temporal-Difference Learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6093230)