Regular policies in abstract dynamic programming
From MaRDI portal
Abstract: We consider challenging dynamic programming models where the associated Bellman equation, and the value and policy iteration algorithms commonly exhibit complex and even pathological behavior. Our analysis is based on the new notion of regular policies. These are policies that are well-behaved with respect to value and policy iteration, and are patterned after proper policies, which are central in the theory of stochastic shortest path problems. We show that the optimal cost function over regular policies may have favorable value and policy iteration properties, which the optimal cost function over all policies need not have. We accordingly develop a unifying methodology to address long standing analytical and algorithmic issues in broad classes of undiscounted models, including stochastic and minimax shortest path problems, as well as positive cost, negative cost, risk-sensitive, and multiplicative cost problems.
Recommendations
Cites work
- A mixed value and policy iteration method for stochastic control with universally measurable policies
- Abstract dynamic programming
- An Analysis of Stochastic Shortest Path Problems
- Contraction Mappings in the Theory Underlying Dynamic Programming
- Dynamic programming and optimal control. Vol. 1.
- Dynamic programming and optimal control. Vol. 2
- Finite state Markovian decision processes
- scientific article; zbMATH DE number 3889341 (Why is no real title available?)
- scientific article; zbMATH DE number 4061056 (Why is no real title available?)
- scientific article; zbMATH DE number 3724212 (Why is no real title available?)
- scientific article; zbMATH DE number 44406 (Why is no real title available?)
- scientific article; zbMATH DE number 51132 (Why is no real title available?)
- scientific article; zbMATH DE number 3623330 (Why is no real title available?)
- scientific article; zbMATH DE number 1325008 (Why is no real title available?)
- scientific article; zbMATH DE number 700091 (Why is no real title available?)
- scientific article; zbMATH DE number 1786118 (Why is no real title available?)
- scientific article; zbMATH DE number 3793773 (Why is no real title available?)
- scientific article; zbMATH DE number 3249162 (Why is no real title available?)
- scientific article; zbMATH DE number 3301983 (Why is no real title available?)
- Monotone Mappings with Application in Dynamic Programming
- Negative Dynamic Programming
- On convergence of value iteration for a class of total cost Markov decision processes
- On terminating Markov decision processes with a risk-averse objective function
- Risk-averse control of undiscounted transient Markov models
- Robust shortest path planning and semicontractive dynamic programming
- Stable Optimal Control and Semicontractive Dynamic Programming
- Stochastic optimal control. The discrete time case
- Stochastic Shortest Path Games
Cited in
(4)
This page was built for publication: Regular policies in abstract dynamic programming
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5348471)