Two Time-Scale Stochastic Approximation with Controlled Markov Noise and Off-Policy Temporal-Difference Learning
From MaRDI portal
Abstract: We present for the first time an asymptotic convergence analysis of two time-scale stochastic approximation driven by `controlled' Markov noise. In particular, both the faster and slower recursions have non-additive controlled Markov noise components in addition to martingale difference noise. We analyze the asymptotic behavior of our framework by relating it to limiting differential inclusions in both time-scales that are defined in terms of the ergodic occupation measures associated with the controlled Markov processes. Finally, we present a solution to the off-policy convergence problem for temporal difference learning with linear function approximation, using our results.
Recommendations
- Stability of Stochastic Approximations With “Controlled Markov” Noise and Temporal Difference Learning
- Policy learning for time-bounded reachability in continuous-time Markov decision processes via doubly-stochastic gradient ascent
- Temporal difference-based policy iteration for optimal control of stochastic systems
- Learning control of finite Markov chains with an explicit trade-off between estimation and control
- A Two-Timescale Stochastic Algorithm Framework for Bilevel Optimization: Complexity Analysis and Application to Actor-Critic
- Risk-averse learning by temporal difference methods with Markov risk measures
- scientific article; zbMATH DE number 7307478
- A Simultaneous Perturbation Stochastic Approximation-Based Actor–Critic Algorithm for Markov Decision Processes
- Approximate Q Learning for Controlled Diffusion Processes and Its Near Optimality
- Simple and optimal methods for stochastic variational inequalities. II: Markovian noise and policy evaluation in reinforcement learning
Cites work
- Applications of a Kushner and Clark lemma to general classes of stochastic algorithms
- Basis function adaptation in temporal difference reinforcement learning
- Convergence and convergence rate of stochastic gradient search in the case of multiple and non-isolated extrema
- scientific article; zbMATH DE number 3855514 (Why is no real title available?)
- scientific article; zbMATH DE number 48727 (Why is no real title available?)
- scientific article; zbMATH DE number 825585 (Why is no real title available?)
- scientific article; zbMATH DE number 1405930 (Why is no real title available?)
- scientific article; zbMATH DE number 3238721 (Why is no real title available?)
- Least squares temporal difference methods: An analysis under general conditions
- Linear stochastic approximation driven by slowly varying Markov chains
- OnActor-Critic Algorithms
- Stochastic approximation with `controlled Markov' noise
- Stochastic approximation with two time scales
- Stochastic Approximations and Differential Inclusions
- Stochastic approximations for finite-state Markov chains
- Weak convergence properties of constrained emphatic temporal-difference learning with constant and slowly diminishing stepsize
Cited in
(14)- Convergence rate of linear two-time-scale stochastic approximation.
- Finite-sample analysis of nonlinear stochastic approximation with applications in reinforcement learning
- Whittle index based Q-learning for restless bandits with average reward
- Linear stochastic approximation driven by slowly varying Markov chains
- Weak convergence properties of constrained emphatic temporal-difference learning with constant and slowly diminishing stepsize
- Stochastic Recursive Inclusions in Two Timescales with Nonadditive Iterate-Dependent Markov Noise
- Proximal gradient temporal difference learning: stable reinforcement learning with polynomial sample complexity
- Finite-time analysis and restarting scheme for linear two-time-scale stochastic approximation
- Simple and optimal methods for stochastic variational inequalities. II: Markovian noise and policy evaluation in reinforcement learning
- A Two-Timescale Stochastic Algorithm Framework for Bilevel Optimization: Complexity Analysis and Application to Actor-Critic
- A Two-Time-Scale Stochastic Optimization Framework with Applications in Control and Reinforcement Learning
- Stochastic approximation and reinforcement learning: the interface and a little beyond
- Stochastic approximation with two time scales: the general case
- The ODE method for asymptotic statistics in stochastic approximation and reinforcement learning
This page was built for publication: Two Time-Scale Stochastic Approximation with Controlled Markov Noise and Off-Policy Temporal-Difference Learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5219302)