A unified algorithm framework for mean-variance optimization in discounted Markov decision processes
From MaRDI portal
(Redirected from Publication:6096629)
Abstract: This paper studies the risk-averse mean-variance optimization in infinite-horizon discounted Markov decision processes (MDPs). The involved variance metric concerns reward variability during the whole process, and future deviations are discounted to their present values. This discounted mean-variance optimization yields a reward function dependent on a discounted mean, and this dependency renders traditional dynamic programming methods inapplicable since it suppresses a crucial property -- time consistency. To deal with this unorthodox problem, we introduce a pseudo mean to transform the untreatable MDP to a standard one with a redefined reward function in standard form and derive a discounted mean-variance performance difference formula. With the pseudo mean, we propose a unified algorithm framework with a bilevel optimization structure for the discounted mean-variance optimization. The framework unifies a variety of algorithms for several variance-related problems including, but not limited to, risk-averse variance and mean-variance optimizations in discounted and average MDPs. Furthermore, the convergence analyses missing from the literature can be complemented with the proposed framework as well. Taking the value iteration as an example, we develop a discounted mean-variance value iteration algorithm and prove its convergence to a local optimum with the aid of a Bellman local-optimality equation. Finally, we conduct a numerical experiment on portfolio management to validate the proposed algorithm.
Cites work
- A mean-variance optimization problem for discounted Markov decision processes
- A possibilistic mean-semivariance-entropy model for multi-period portfolio selection with transaction costs
- Advances in prospect theory: cumulative representation of uncertainty
- Analysis and improvement of policy gradient estimation
- Continuous-time mean-variance portfolio selection: a stochastic LQ framework
- scientific article; zbMATH DE number 51708 (Why is no real title available?)
- scientific article; zbMATH DE number 5685899 (Why is no real title available?)
- Markowitz's Mean-Variance Portfolio Selection with Regime Switching: A Continuous-Time Model
- Mean-variance analysis of option contracts in a two-echelon supply chain
- Mean-variance optimization of discrete time discounted Markov decision processes
- Mean-Variance Tradeoffs in an Undiscounted MDP
- Mean-Variance Tradeoffs in an Undiscounted MDP: The Unichain Case
- Multilevel optimization modeling for risk-averse stochastic programming
- Optimal dynamic portfolio selection: multiperiod mean-variance formulation
- Optimization of Markov decision processes under the variance criterion
- Portfolio Optimization with Nonparametric Value at Risk: A Block Coordinate Descent Method
- Reinforcement learning. An introduction
- Sample-Path Optimality and Variance-Minimization of Average Cost Markov Control Processes
- Sensitivity Analysis for Mean-Variance Portfolio Problems
- Stochastic learning and optimization. A sensitivity-based approach.
- Survey on multi-period mean-variance portfolio selection model
- The variance of discounted Markov decision processes
- Variance-Penalized Markov Decision Processes
- Variance-penalized Markov decision processes: dynamic programming and reinforcement learning techniques
Cited in
(4)- Variance minimization for constrained discounted continuous-time MDPs with exponentially distributed stopping times
- Mean-Variance Tradeoffs in an Undiscounted MDP
- The setting of risk prevention threshold of the perishable product supply chain with retailer's uncertain risk preference
- Independent policy gradient-based reinforcement learning for economic and reliable energy management of multi-microgrid systems
This page was built for publication: A unified algorithm framework for mean-variance optimization in discounted Markov decision processes
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6096629)