On the policy improvement algorithm in continuous time
From MaRDI portal
Abstract: We develop a general approach to the Policy Improvement Algorithm (PIA) for stochastic control problems for continuous-time processes. The main results assume only that the controls lie in a compact metric space and give general sufficient conditions for the PIA to be well-defined and converge in continuous time (i.e. without time discretisation). It emerges that the natural context for the PIA in continuous time is weak stochastic control. We give examples of control problems demonstrating the need for the weak formulation as well as diffusion-based classes of problems where the PIA in continuous time is applicable.
Recommendations
- The policy iteration algorithm for average continuous control of piecewise deterministic Markov processes
- Policy iteration for continuous-time average reward Markov decision processes in Polish spaces
- Policy iterations for reinforcement learning problems in continuous time and space -- fundamental theory and methods
- Exponential convergence and stability of Howard's policy improvement algorithm for controlled diffusions
- Policy gradient in continuous time
Cited in
(18)- A policy iteration algorithm for the American put option and free boundary control problems
- Extensions of the deep Galerkin method
- A modified MSA for stochastic control problems
- Policy iterations for reinforcement learning problems in continuous time and space -- fundamental theory and methods
- Policy gradient in continuous time
- A Class of Decision Processes Showing Policy-Improvement/Newton–Raphson Equivalence
- scientific article; zbMATH DE number 6982305 (Why is no real title available?)
- On the parabolic equation for portfolio problems
- Exponential convergence and stability of Howard's policy improvement algorithm for controlled diffusions
- A Policy Improvement Method in Constrained Stochastic Dynamic Programming
- The policy iteration algorithm for average continuous control of piecewise deterministic Markov processes
- On continuous time agents
- The modified MSA, a gradient flow and convergence
- Policy iteration for exploratory Hamilton-Jacobi-Bellman equations
- Convergence of policy iteration for entropy-regularized stochastic control problems
- Optimal control by policy iterations and constrained Gaussian process regressions
- Convergence analysis for entropy-regularized control problems: a probabilistic approach
- The Howard's policy iteration and convergence for optimal dividend under compound-Poisson model
This page was built for publication: On the policy improvement algorithm in continuous time
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2974868)