Exponential convergence and stability of Howard's policy improvement algorithm for controlled diffusions
From MaRDI portal
(Redirected from Publication:5111071)
Abstract: Optimal control problems are inherently hard to solve as the optimization must be performed simultaneously with updating the underlying system. Starting from an initial guess, Howard's policy improvement algorithm separates the step of updating the trajectory of the dynamical system from the optimization and iterations of this should converge to the optimal control. In the discrete space-time setting this is often the case and even rates of convergence are known. In the continuous space-time setting of controlled diffusion the algorithm consists of solving a linear PDE followed by maximization problem. This has been shown to converge, in some situations, however no global rate of is known. The first main contribution of this paper is to establish global rate of convergence for the policy improvement algorithm and a variant, called here the gradient iteration algorithm. The second main contribution is the proof of stability of the algorithms under perturbations to both the accuracy of the linear PDE solution and the accuracy of the maximization step. The proof technique is new in this context as it uses the theory of backward stochastic differential equations.
Recommendations
- On the policy improvement algorithm for ergodic risk-sensitive control
- Policy iteration algorithm for singular controlled diffusion processes
- On the policy improvement algorithm in continuous time
- Rates of convergence for the policy iteration method for mean field games systems
- Some Convergence Results for Howard's Algorithm
Cites work
- scientific article; zbMATH DE number 3126094 (Why is no real title available?)
- scientific article; zbMATH DE number 3148886 (Why is no real title available?)
- Average Optimality in Markov Control Processes via Discounted-Cost Problems and Linear Programming
- Continuous-time stochastic control and optimization with financial applications
- Control improvement for jump-diffusion processes with applications to finance
- Controlled Markov processes and viscosity solutions
- Convergence Properties of Policy Iteration
- FUNCTIONAL EQUATIONS IN THE THEORY OF DYNAMIC PROGRAMMING. V. POSITIVITY AND QUASI-LINEARITY
- Infinite horizon backward stochastic differential equations and elliptic equations in Hilbert spaces.
- Markovian quadratic and superquadratic BSDEs with an unbounded terminal condition
- On finite-difference approximations for normalized Bellman equations
- On the Convergence of Policy Iteration in Stationary Dynamic Programming
- On the convergence of policy iteration for controlled diffusions
- On the policy improvement algorithm in continuous time
- SOME NEW RESULTS IN THE THEORY OF CONTROLLED DIFFUSION PROCESSES
- Some Convergence Results for Howard's Algorithm
- The rate of convergence of finite-difference approximations for parabolic bellman equations with Lipschitz coefficients in cylindrical domains
Cited in
(23)- A policy gradient framework for stochastic optimal control problems with global convergence guarantee
- Rates of convergence for the policy iteration method for mean field games systems
- On the policy improvement algorithm for ergodic risk-sensitive control
- Exploratory LQG mean field games with entropy regularization
- Gradient flows for regularized stochastic control problems
- Policy iteration for the deterministic control problems -- a viscosity approach
- A modified MSA for stochastic control problems
- Policy iteration for exploratory Hamilton-Jacobi-Bellman equations
- On the policy improvement algorithm in continuous time
- Linear Convergence of a Policy Gradient Method for Some Finite Horizon Continuous Time Control Problems
- Convergence of policy iteration for entropy-regularized stochastic control problems
- A neural network-based policy iteration algorithm with global \(H^2\)-superlinear convergence for stochastic games on domains
- Convergence analysis for entropy-regularized control problems: a probabilistic approach
- Policy iteration for nonconvex viscous Hamilton-Jacobi equations
- The modified MSA, a gradient flow and convergence
- Market based mechanisms for incentivising exchange liquidity provision
- A policy iteration method for mean field games
- Reinforcement Learning for Linear-Convex Models with Jumps via Stability Analysis of Feedback Controls
- A modified method of successive approximations for stochastic recursive optimal control problems
- Improved order 1/4 convergence for piecewise constant policy approximation of stochastic control problems
- Policy iteration method for time-dependent mean field games systems with non-separable Hamiltonians
- Mirror descent for stochastic control problems with measure-valued controls
- Policy iteration algorithm for singular controlled diffusion processes
This page was built for publication: Exponential convergence and stability of Howard's policy improvement algorithm for controlled diffusions
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5111071)