Empirical dynamic programming
dynamic programmingempirical methodsMarkov decision processesprobabilistic fixed pointsrandom operatorssimulation
Random dynamical systems (37H99) Simulation of dynamical systems (37M05) Random linear operators (47B80) Dynamic programming in optimal control and differential games (49L20) Random operators and equations (aspects of stochastic analysis) (60H25) Empirical decision procedures; empirical Bayes procedures (62C12) Numerical mathematical programming methods (65K05) Stochastic programming (90C15) Dynamic programming (90C39) Markov and semi-Markov decision processes (90C40) Optimal stochastic control (93E20)
- \({\mathcal Q}\)-learning
- 10.1162/153244303768966102
- A Stochastic Approximation Method
- A survey of some simulation-based algorithms for Markov decision processes
- Actor-Critic--Type Learning Algorithms for Markov Decision Processes
- Analysis of recursive stochastic algorithms
- Approximate Fixed Point Iteration with an Application to Infinite Horizon Markov Decision Processes
- Approximate policy iteration: a survey and some new methods
- Approximations of Dynamic Programs, I
- Approximations of Dynamic Programs, II
- Associative search network: A reinforcement learning associative memory
- Comparison methods for stochastic models and risks
- CONVERGENCE OF SIMULATION-BASED POLICY ITERATION
- Convergence rate of linear two-time-scale stochastic approximation.
- Finite-time bounds for fitted value iteration
- Functional Approximations and Dynamic Programming
- scientific article; zbMATH DE number 5957196 (Why is no real title available?)
- scientific article; zbMATH DE number 3361677 (Why is no real title available?)
- Learning algorithms for Markov decision processes with average cost
- Neural Network Learning
- Performance guarantees for empirical Markov decision processes with applications to multiperiod inventory models
- Q-Learning for Risk-Sensitive Control
- Simulation-based optimization of Markov decision processes: an empirical process theory approach
- Simulation‐based Uniform Value Function Estimates of Markov Decision Processes
- Stochastic Estimation of the Maximum of a Regression Function
- Stochastic Games
- The Complexity of Markov Decision Processes
- The O.D.E. Method for Convergence of Stochastic Approximation and Reinforcement Learning
- Using Randomization to Break the Curse of Dimensionality
- Stochastic and adaptive optimal control of uncertain interconnected systems: a data-driven approach
- Gradient-bounded dynamic programming for submodular and concave extensible value functions with probabilistic performance guarantees
- Robustness to incorrect models and data-driven learning in average-cost optimal stochastic control
- A concentration bound for contractive stochastic approximation
- A simulation-based approach to stochastic dynamic programming
- scientific article; zbMATH DE number 1195626 (Why is no real title available?)
- Convergence of Recursive Stochastic Algorithms Using Wasserstein Divergence
- Mean-field controls with Q-learning for cooperative MARL: convergence and complexity analysis
- Some limit properties of Markov chains induced by recursive stochastic algorithms
- Distributionally robust optimization for sequential decision-making
- Dynamic policy programming
- Empirical Q-value iteration
- Anderson acceleration for partially observable Markov decision processes: a maximum entropy approach
- Error analysis for approximate CVaR-optimal control with a maximum cost
- Dynamic programs on partially ordered sets
This page was built for publication: Empirical dynamic programming
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2806811)