Q() with off-policy corrections
From MaRDI portal
Publication:2831390
Abstract: We propose and analyze an alternate approach to off-policy multi-step temporal difference learning, in which off-policy returns are corrected with the current Q-function in terms of rewards, rather than with the target policy in terms of transition probabilities. We prove that such approximate corrections are sufficient for off-policy convergence both in policy evaluation and control, provided certain conditions. These conditions relate the distance between the target and behavior policies, the eligibility trace parameter and the discount factor, and formalize an underlying tradeoff in off-policy TD(). We illustrate this theoretical relationship empirically on a continuous-state control task.
Recommendations
Cites work
- \({\mathcal Q}\)-learning
- \(\text{Q}(\lambda)\) with off-policy corrections
- Analytical mean squared error curves for temporal difference learning
- scientific article; zbMATH DE number 3126094 (Why is no real title available?)
- scientific article; zbMATH DE number 1321699 (Why is no real title available?)
- scientific article; zbMATH DE number 700091 (Why is no real title available?)
Cited in
(15)- Off-policy temporal difference learning with distribution adaptation in fast mixing chains
- Classification with costly features as a sequential decision-making problem
- TD-regularized actor-critic methods
- An emphatic approach to the problem of off-policy temporal-difference learning
- \(\text{Q}(\lambda)\) with off-policy corrections
- Off-policy learning with eligibility traces: a survey
- Off-policy linear temporal difference learning algorithms with a generalized oblique projection
- Reinforcement learning in sparse-reward environments with hindsight policy gradients
- Deep Reinforcement Learning: A State-of-the-Art Walkthrough
- Double reinforcement learning for efficient off-policy evaluation in Markov decision processes
- Deep exploration via randomized value functions
- Optimistic reinforcement learning by forward Kullback-Leibler divergence optimization
- Gradient temporal-difference learning for off-policy evaluation using emphatic weightings
- An approximate policy iteration viewpoint of actor-critic algorithms
- Concentration of contractive stochastic approximation: additive and multiplicative noise
This page was built for publication: \(\text{Q}(\lambda)\) with off-policy corrections
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2831390)