Value-Gradient Based Formulation of Optimal Control Problem and Machine Learning Algorithm
From MaRDI portal
Abstract: Optimal control problem is typically solved by first finding the value function through Hamilton-Jacobi equation (HJE) and then taking the minimizer of the Hamiltonian to obtain the control. In this work, instead of focusing on the value function, we propose a new formulation for the gradient of the value function (value-gradient) as a decoupled system of partial differential equations in the context of continuous-time deterministic discounted optimal control problem. We develop an efficient iterative scheme for this system of equations in parallel by utilizing the properties that they share the same characteristic curves as the HJE for the value function. For the theoretical part, we prove that this iterative scheme converges linearly in sense for some suitable exponent in a weight function. For the numerical method, we combine characteristic line method with machine learning techniques. Specifically, we generate multiple characteristic curves at each policy iteration from an ensemble of initial states, and compute both the value function and its gradient simultaneously on each curve as the labelled data. Then supervised machine learning is applied to minimize the weighted squared loss for both the value function and its gradients. Experimental results demonstrate that this new method not only significantly increases the accuracy but also improves the efficiency and robustness of the numerical estimates, particularly with less amount of characteristics data or fewer training steps.
Recommendations
- Stochastic optimal control via forward and backward stochastic differential equations and importance sampling
- A projected primal-dual gradient optimal control method for deep reinforcement learning
- Solving stochastic optimal control problem via stochastic maximum principle with deep learning method
- Data-Driven Tensor Train Gradient Cross Approximation for Hamilton–Jacobi–Bellman Equations
- A mean-field optimal control formulation of deep learning
Cites work
- A splitting method for overcoming the curse of dimensionality in Hamilton-Jacobi equations arising from nonlinear optimal control and differential games with applications to trajectory generation
- Adaptive deep learning for high-dimensional Hamilton-Jacobi-Bellman equations
- Algorithm for Hamilton-Jacobi equations in density space via a generalized Hopf formula
- Algorithm for overcoming the curse of dimensionality for certain non-convex Hamilton-Jacobi equations, projections and differential games
- Algorithm for overcoming the curse of dimensionality for time-dependent non-convex Hamilton-Jacobi equations arising from optimal control and differential games problems
- Algorithms for overcoming the curse of dimensionality for certain Hamilton-Jacobi equations arising in control theory and elsewhere
- Algorithms for solving high dimensional PDEs: from nonlinear Monte Carlo to machine learning
- An efficient policy iteration algorithm for dynamic programming equations
- Approximate solutions to the time-invariant Hamilton-Jacobi-Bellman equation
- Controlled Markov processes and viscosity solutions
- Deep neural networks algorithms for stochastic control problems on finite horizon: convergence analysis
- Deep neural networks algorithms for stochastic control problems on finite horizon: numerical applications
- Dynamic programming and optimal control. Vol. 2.
- Estimation and control of dynamical systems
- Fast Sweeping Algorithms for a Class of Hamilton--Jacobi Equations
- Fronts propagating with curvature-dependent speed: Algorithms based on Hamilton-Jacobi formulations
- Galerkin approximations of the generalized Hamilton-Jacobi-Bellman equation
- scientific article; zbMATH DE number 3126094 (Why is no real title available?)
- scientific article; zbMATH DE number 3128787 (Why is no real title available?)
- scientific article; zbMATH DE number 3148886 (Why is no real title available?)
- scientific article; zbMATH DE number 140583 (Why is no real title available?)
- scientific article; zbMATH DE number 3505708 (Why is no real title available?)
- scientific article; zbMATH DE number 7626721 (Why is no real title available?)
- scientific article; zbMATH DE number 5681750 (Why is no real title available?)
- Machine Learning and Control Theory
- Mitigating the curse of dimensionality: sparse grid characteristics method for optimal feedback control and HJB equations
- Neural network architectures using min-plus algebra for solving certain high-dimensional optimal control problems and Hamilton-Jacobi PDEs
- On some neural network architectures that can represent viscosity solutions of certain high dimensional Hamilton-Jacobi partial differential equations
- On the Convergence of Policy Iteration in Stationary Dynamic Programming
- Overcoming the curse of dimensionality for some Hamilton-Jacobi partial differential equations via neural network architectures
- Polynomial approximation of high-dimensional Hamilton-Jacobi-Bellman equations and applications to feedback control of semilinear parabolic PDEs
- Reinforcement learning. An introduction
- Semi-Lagrangian approximation schemes for linear and Hamilton-Jacobi equations
- Solving high-dimensional partial differential equations using deep learning
- Successive Galerkin approximation algorithms for nonlinear optimal and robust control
Cited in
(3)
This page was built for publication: Value-Gradient Based Formulation of Optimal Control Problem and Machine Learning Algorithm
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6040292)