The O.D.E. Method for Convergence of Stochastic Approximation and Reinforcement Learning
From MaRDI portal
Recommendations
- On the convergence of reinforcement learning
- Concentration of Contractive Stochastic Approximation and Reinforcement Learning
- 10.1162/153244303768966102
- Stochastic Approximation for Nonexpansive Maps: Application to Q-Learning Algorithms
- Approximate gradient methods in policy-space optimization of Markov reward processes
- Finite-sample analysis of nonlinear stochastic approximation with applications in reinforcement learning
- Convergence results for single-step on-policy reinforcement-learning algorithms
Cited in
(95)- Cooperative dynamics and Wardrop equilibria
- Natural actor-critic algorithms
- Asynchronous stochastic approximation and Q-learning
- Reinforcement learning for long-run average cost.
- Stability of annealing schemes and related processes
- An online prediction algorithm for reinforcement learning with linear function approximation using cross entropy method
- Variance-constrained actor-critic algorithms for discounted and average reward MDPs
- Asymptotic bias of stochastic gradient search
- Approachability in Stackelberg stochastic games with vector costs
- Q-learning for Markov decision processes with a satisfiability criterion
- Popularity signals in trial-offer markets with social influence and position bias
- Error bounds for constant step-size \(Q\)-learning
- Convergence and convergence rate of stochastic gradient search in the case of multiple and non-isolated extrema
- On the convergence of stochastic approximations under a subgeometric ergodic Markov dynamic
- A stochastic primal-dual method for optimization with conditional value at risk constraints
- Concentration bounds for temporal difference learning with linear function approximation: the case of batch data and uniform sampling
- Revisiting the ODE method for recursive algorithms: fast convergence using quasi stochastic approximation
- An ODE method to prove the geometric convergence of adaptive stochastic algorithms
- What may lie ahead in reinforcement learning
- Fundamental design principles for reinforcement learning algorithms
- Finite-sample analysis of nonlinear stochastic approximation with applications in reinforcement learning
- A sojourn-based approach to semi-Markov reinforcement learning
- Simultaneous perturbation Newton algorithms for simulation optimization
- Non-asymptotic error bounds for constant stepsize stochastic approximation for tracking mobile agents
- An information-theoretic analysis of return maximization in reinforcement learning
- Online calibrated forecasts: memory efficiency versus universality for learning in games
- An adaptive optimization scheme with satisfactory transient performance
- A stability criterion for two timescale stochastic approximation schemes
- Avoidance of traps in stochastic approximation
- Linear stochastic approximation driven by slowly varying Markov chains
- Boundedness of iterates in \(Q\)-learning
- Multi-armed bandits based on a variant of simulated annealing
- Event-driven stochastic approximation
- Reinforcement learning based algorithms for average cost Markov decision processes
- Two-timescale stochastic gradient descent in continuous time with applications to joint online parameter estimation and optimal sensor placement
- Empirical dynamic programming
- Nonlinear gossip
- Multiscale Q-learning with linear function approximation
- Deceptive Reinforcement Learning Under Adversarial Manipulations on Cost Signals
- Learning to control a structured-prediction decoder for detection of HTTP-layer DDoS attackers
- Stochastic Recursive Inclusions in Two Timescales with Nonadditive Iterate-Dependent Markov Noise
- Oja's algorithm for graph clustering, Markov spectral decomposition, and risk sensitive control
- Stochastic approximation with long range dependent and heavy tailed noise
- An online actor-critic algorithm with function approximation for constrained Markov decision processes
- On stochastic gradient and subgradient methods with adaptive steplength sequences
- Stabilization of stochastic approximation by step size adaptation
- Stochastic Approximation for Nonexpansive Maps: Application to Q-Learning Algorithms
- Distributed stochastic approximation with local projections
- Risk-averse learning by temporal difference methods with Markov risk measures
- Finite-time performance of distributed temporal-difference learning with linear function approximation
- A finite time analysis of temporal difference learning with linear function approximation
- Convergence of Recursive Stochastic Algorithms Using Wasserstein Divergence
- Iterative learning control using faded measurements without system information: a gradient estimation approach
- Some limit properties of Markov chains induced by recursive stochastic algorithms
- A Diffusion Approximation Theory of Momentum Stochastic Gradient Descent in Nonconvex Optimization
- Stochastic recursive inclusions with non-additive iterate-dependent Markov noise
- Model-Free Reinforcement Learning for Stochastic Parity Games
- Risk-Sensitive Reinforcement Learning via Policy Gradient Search
- Technical note: Consistency analysis of sequential learning under approximate Bayesian inference
- Full gradient DQN reinforcement learning: a provably convergent scheme
- Is Temporal Difference Learning Optimal? An Instance-Dependent Analysis
- Asynchronous stochastic approximation with differential inclusions
- Quasi-Newton smoothed functional algorithms for unconstrained and constrained simulation optimization
- The Borkar-Meyn theorem for asynchronous stochastic approximations
- Analyzing approximate value iteration algorithms
- Concentration of Contractive Stochastic Approximation and Reinforcement Learning
- Accelerated and Instance-Optimal Policy Evaluation with Linear Function Approximation
- A sensitivity formula for risk-sensitive cost and the actor-critic algorithm
- A Small Gain Analysis of Single Timescale Actor Critic
- Convergence of stochastic approximation via martingale and converse Lyapunov methods
- A Discrete-Time Switching System Analysis of Q-Learning
- On the sample complexity of actor-critic method for reinforcement learning with function approximation
- Gradient temporal-difference learning for off-policy evaluation using emphatic weightings
- Target Network and Truncation Overcome the Deadly Triad in \(\boldsymbol{Q}\)-Learning
- Multi-agent natural actor-critic reinforcement learning algorithms
- An actor-critic algorithm with function approximation for discounted cost constrained Markov decision processes
- A Two-Time-Scale Stochastic Optimization Framework with Applications in Control and Reinforcement Learning
- Finite-time error bounds for distributed linear stochastic approximation
- Distributed non-linear robust consensus-based sensor calibration for networked control systems
- Convergence of stochastic approximation via martingale and converse Lyapunov methods
- Convergence rates for stochastic approximation: biased noise with unbounded variance, and applications
- Analysis of multiscale reinforcement Q-learning algorithms for mean field control games
- Stochastic approximation and reinforcement learning: the interface and a little beyond
- Error analysis for approximate CVaR-optimal control with a maximum cost
- The ODE method for asymptotic statistics in stochastic approximation and reinforcement learning
- The ODE method for stochastic approximation and reinforcement learning with Markovian noise
- Stochastic approximation in infinite dimensions
- Markovian foundations for quasi-stochastic approximation
- Charge-based control of DiffServ-like queues
- Recent advances in stochastic approximation with applications to optimization and fixed point problems
- Asynchronous stochastic approximation with applications to average-reward reinforcement learning
- Accuracy of discretely sampled stochastic policies in continuous-time reinforcement learning
- On the analysis of model-free methods for the linear quadratic regulator
- Distributed proximal algorithms for constrained resource allocation over continuous and discrete communication
- A new learning algorithm for optimal stopping
This page was built for publication: The O.D.E. Method for Convergence of Stochastic Approximation and Reinforcement Learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4943730)