Stable Optimal Control and Semicontractive Dynamic Programming
From MaRDI portal
Abstract: We consider discrete-time infinite horizon deterministic optimal control problems with nonnegative cost per stage, and a destination that is cost-free and absorbing. The classical linear-quadratic regulator problem is a special case. Our assumptions are very general, and allow the possibility that the optimal policy may not be stabilizing the system, e.g., may not reach the destination either asymptotically or in a finite number of steps. We introduce a new unifying notion of stable feedback policy, based on perturbation of the cost per stage, which in addition to implying convergence of the generated states to the destination, quantifies the speed of convergence. We consider the properties of two distinct cost functions: , the overall optimal, and , the restricted optimal over just the stable policies. Different classes of stable policies (with different speeds of convergence) may yield different values of . We show that for any class of stable policies, is a solution of Bellman's equation, and we characterize the smallest and the largest solutions: they are , and , the restricted optimal cost function over the class of (finitely) terminating policies. We also characterize the regions of convergence of various modified versions of value and policy iteration algorithms, as substitutes for the standard algorithms, which may not work in general.
Recommendations
- Optimal semistable control for continuous-time linear systems
- Stabilizability in optimal control
- Dynamic consistency for stochastic optimal control problems
- Structural stability of optimal control problems
- Optimal control for linear and nonlinear semistabilization
- scientific article; zbMATH DE number 4004051
- scientific article; zbMATH DE number 1400162
- Stochastic optimal control of state constrained systems
- Stochastic optimal control and linear programming approach
- Optimal control of a setvalued stochastic dynamic system
Cites work
- Affine Monotonic and Risk-Sensitive Models in Dynamic Programming
- Dynamic programming and optimal control. Vol. 1.
- Dynamic programming and optimal control. Vol. 2
- scientific article; zbMATH DE number 1321699 (Why is no real title available?)
- scientific article; zbMATH DE number 700091 (Why is no real title available?)
- scientific article; zbMATH DE number 3439537 (Why is no real title available?)
- scientific article; zbMATH DE number 802915 (Why is no real title available?)
- scientific article; zbMATH DE number 3388902 (Why is no real title available?)
- Negative Dynamic Programming
- Optimal adaptive control and differential games by reinforcement learning principles
- Regular policies in abstract dynamic programming
- Robust shortest path planning and semicontractive dynamic programming
- Stochastic optimal control. The discrete time case
Cited in
(16)- Computing efficient steady state policies for deterministic dynamic programs. I
- Drift counteraction optimal control for deterministic systems and enhancing convergence of value iteration
- Improved value iteration for neural-network-based stochastic optimal control design
- Stable synthesis of optimal control in stationary extremal problems
- Stable and robust LQR design via scenario approach
- Logarithmic regret in online linear quadratic control using Riccati updates
- A mixed value and policy iteration method for stochastic control with universally measurable policies
- A semi-Lagrangian algorithm in policy space for hybrid optimal control problems
- Temporal difference-based policy iteration for optimal control of stochastic systems
- Simple and optimal methods for stochastic variational inequalities. II: Markovian noise and policy evaluation in reinforcement learning
- Regular policies in abstract dynamic programming
- On the relation between dynamic regret and closed-loop stability
- Accessibility and stabilization by infinite horizon optimal control with negative discounting
- Optimal control of linear cost networks
- Learning near-optimal broadcasting intervals in decentralized multi-agent systems using online least-square policy iteration
- Generalized value iteration for discounted optimal control with stability analysis
This page was built for publication: Stable Optimal Control and Semicontractive Dynamic Programming
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3130440)