Quantile Markov Decision Processes
From MaRDI portal
Abstract: The goal of a traditional Markov decision process (MDP) is to maximize expected cumulativereward over a defined horizon (possibly infinite). In many applications, however, a decision maker may beinterested in optimizing a specific quantile of the cumulative reward instead of its expectation. In this paperwe consider the problem of optimizing the quantiles of the cumulative rewards of a Markov decision process(MDP), which we refer to as a quantile Markov decision process (QMDP). We provide analytical resultscharacterizing the optimal QMDP value function and present a dynamic programming-based algorithm tosolve for the optimal policy. The algorithm also extends to the MDP problem with a conditional value-at-risk(CVaR) objective. We illustrate the practical relevance of our model by evaluating it on an HIV treatmentinitiation problem, where patients aim to balance the potential benefits and risks of the treatment.
Recommendations
- Computing quantiles in Markov reward models
- Qualitative analysis of partially-observable Markov decision processes
- Markov decision processes with quasi-hyperbolic discounting
- Markov decision processes
- scientific article; zbMATH DE number 4154239
- scientific article; zbMATH DE number 420890
- Markov decision processes
- MARKOV DECISION PROCESSES
- Quantile Maximization in Decision Theory*
Cites work
- Average Cost Semi-Markov Decision Processes and the Control of Queueing Systems
- Computing quantiles in Markov reward models
- Convex Approximations of Chance Constrained Programs
- Dynamic programming in constrained Markov decision processes
- scientific article; zbMATH DE number 1348599 (Why is no real title available?)
- scientific article; zbMATH DE number 1095138 (Why is no real title available?)
- Markov decision problems where means bound variances
- Markov decision processes with average-value-at-risk criteria
- Optimal stopping of Markov processes: Hilbert space theory, approximation algorithms, and an application to pricing high-dimensional financial derivatives
- Optimizing the simultaneous management of blood pressure and cholesterol for type 2 diabetes patients
- Percentile Optimization for Markov Decision Processes with Parameter Uncertainty
- Percentile performance criteria for limiting average Markov decision processes
- Risk neutral and risk averse stochastic dual dynamic programming method
- Risk-averse approximate dynamic programming with quantile-based risk measures
- Risk-averse dynamic programming for Markov decision processes
- Risk-Sensitive Markov Decision Processes
- Robust Control of Markov Decision Processes with Uncertain Transition Matrices
- Robust Markov Decision Processes
- The Optimal Time to Initiate HIV Therapy Under Ordered Health States
- Tight approximations of dynamic risk measures
- Time-consistent decisions and temporal decomposition of coherent risk functionals
Cited in
(10)- Quantile Maximization in Decision Theory*
- Ordinal decision models for Markov decision processes
- scientific article; zbMATH DE number 7297838 (Why is no real title available?)
- Risk-averse approximate dynamic programming with quantile-based risk measures
- Percentile queries in multi-dimensional Markov decision processes
- CVaR-based optimization of environmental flow via the Markov lift of a mixed moving average process
- Learning risk preferences in Markov decision processes: an application to the fourth down decision in the national football league
- Team variance optimization of n-player stochastic games with separately controlled chains
- Sequential decision-making under uncertainty: a robust MDPs review
- Zero-sum risk-sensitive continuous-time stochastic games with unbounded reward and transition rates in Borel spaces
This page was built for publication: Quantile Markov Decision Processes
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5095150)