On the Convergence of Stochastic Iterative Dynamic Programming Algorithms
From MaRDI portal
Recommendations
Cites work
Cited in
(55)- Reinforcement distribution in fuzzy Q-learning
- Stochastic convexity in dynamic programming
- Convergence results for single-step on-policy reinforcement-learning algorithms
- A unified framework for stochastic optimization
- The convergence of \(TD(\lambda)\) for general \(\lambda\)
- Linear least-squares algorithms for temporal difference learning
- On the worst-case analysis of temporal-difference learning algorithms
- Reinforcement learning with replacing eligibility traces
- Error bounds for constant step-size \(Q\)-learning
- Revisiting the ODE method for recursive algorithms: fast convergence using quasi stochastic approximation
- Restricted gradient-descent algorithm for value-function approximation in reinforcement learning
- The asymptotic equipartition property in reinforcement learning and its relation to return maximization
- Adaptive stock trading with dynamic asset allocation using reinforcement learning
- An optimal control approach to mode generation in hybrid systems
- Boundedness of iterates in \(Q\)-learning
- On the convergence of reinforcement learning with Monte Carlo exploring starts
- A simulation-based approach to stochastic dynamic programming
- Approximate policy iteration: a survey and some new methods
- Stochastic adaptation of importance sampler
- Perspectives of approximate dynamic programming
- Q-learning and policy iteration algorithms for stochastic shortest path problems
- Convergence in unconstrained discrete-time differential dynamic programming
- The optimal unbiased value estimator and its relation to LSTD, TD and MC
- Speed of Convergence and Stopping Rules in an Iterative Planning Procedure for Nonconvex Economies
- TD(λ) learning without eligibility traces: a theoretical analysis
- Bayesian exploration for approximate dynamic programming
- Risk-averse learning by temporal difference methods with Markov risk measures
- Some limit properties of Markov chains induced by recursive stochastic algorithms
- Cooperation between independent market makers
- A Q-Learning Algorithm for Discrete-Time Linear-Quadratic Control with Random Parameters of Unknown Distribution: Convergence and Stabilization
- Technical note: Consistency analysis of sequential learning under approximate Bayesian inference
- Deep Reinforcement Learning: A State-of-the-Art Walkthrough
- Full gradient DQN reinforcement learning: a provably convergent scheme
- Adaptive learning algorithm convergence in passive and reactive environments
- Is Temporal Difference Learning Optimal? An Instance-Dependent Analysis
- Risk-averse approximate dynamic programming with quantile-based risk measures
- SOLVING DYNAMIC WILDLIFE RESOURCE OPTIMIZATION PROBLEMS USING REINFORCEMENT LEARNING
- REINFORCEMENT LEARNING WITH GOAL-DIRECTED ELIGIBILITY TRACES
- Empirical Q-value iteration
- Asymptotic analysis of temporal-difference learning algorithms with constant step-sizes
- Asymptotic analysis of temporal-difference learning algorithms with constant step-sizes
- Stochastic approximation algorithms: overview and recent trends.
- Convergence of least squares learning in self-referential discontinuous stochastic models.
- A Discrete-Time Switching System Analysis of Q-Learning
- A novel policy based on action confidence limit to improve exploration efficiency in reinforcement learning
- Approximate Q Learning for Controlled Diffusion Processes and Its Near Optimality
- Target Network and Truncation Overcome the Deadly Triad in \(\boldsymbol{Q}\)-Learning
- A lexicographic optimization approach for a bi-objective parallel-machine scheduling problem minimizing total quality loss and total tardiness
- Stochastic Fixed-Point Iterations for Nonexpansive Maps: Convergence and Error Bounds
- Platform design when sellers use pricing algorithms
- Robotics and artificial intelligence
- Robust Q-learning algorithm for Markov decision processes under Wasserstein uncertainty
- Stochastic approximation and reinforcement learning: the interface and a little beyond
- Stochastic approximation in infinite dimensions
- Reinforcement learning algorithms with function approximation: recent advances and applications
This page was built for publication: On the Convergence of Stochastic Iterative Dynamic Programming Algorithms
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4323346)