Simulation-based algorithms for Markov decision processes.
This book describes algorithms for computing value functions and optimal policies for discounted Markov Decision Processes (MDPs) via simulation. The basic question addressed in this monograph is how to compute optimal policies and/or value function estimates if the expected one-step rewards and transition probabilities can only be simulated rather than given in an explicit form. The monograph consists of five chapters. Chapter 1, Markov Decision Processes, introduces the policy iteration and value iteration algorithms for finite-horizon discounted MDPs and the rolling horizon approach to infinite-horizon discounted MDPs. Chapter 2, Multi-Stage Sampling Algorithms, discusses finite-horizon problems. Given the total number of allowed simulations at a given state, it provides algorithms on how to allocate these simulations in sampling the admissible actions in order to find good policies and/or value function estimates. Chapter 3, Population-based Evolutionary Approaches, is devoted to versions of the policy iteration algorithm, when the sets of available actions are large and policy iteration cannot be conducted in the explicit form. Simulation is used to generate policies for policy improvement. Most of this chapter deals with the situation when transition probabilities and one-step rewards are given explicitly. Then the results are extended to the situation when these values are simulated. Chapter 4, Model Reference Adaptive Search, introduces and describes a global optimization technique named in the title of this chapter and abbreviated as MRAS. MRAS is first introduced in this chapter for deterministic optimization problems, then for stochastic optimization problems, and then for MDPs. When used for solving MDPs, MRAS works by randomly generating admissible policies from some distribution on the policy space. Then the distribution is updated and policies are generated again. The sequence of distributions should converge to a distribution concentrated at a point. This point is an optimal solution. As the authors indicate, MRAS is closely related to the cross-entropy method. Chapter 5, On-line Control Methods via Simulation, deals with applications of the rolling horizon approach to action selection in real time. This chapter describes and investigates approximate rolling horizon algorithms based on simulation.
- Approximation of discounted minimax Markov control problems and zero-sum Markov games using Hausdorff and Wasserstein distances
- Coupling based estimation approaches for the average reward performance potential in Markov chains
- Approximate stochastic annealing for online control of infinite horizon Markov decision processes
- Simulation-based algorithms for Markov decision processes
- Policy-based branch-and-bound for infinite-horizon multi-model Markov decision processes
- Stochastic approximations of constrained discounted Markov decision processes
- Optimization of Markov decision processes under the variance criterion
- Mean field Markov decision processes
- Computing optimal policies for Markovian decision processes using simulation
- Planning with Markov decision processes. An AI perspective
- Approximate policy iteration: a survey and some new methods
- A review of stochastic algorithms with continuous value function approximation and some new approximate policy iteration algorithms for multidimensional continuous applications
- An evolutionary random policy search algorithm for solving Markov decision processes
- CONIC TRADING IN A MARKOVIAN STEADY STATE
- Solving average cost Markov decision processes by means of a two-phase time aggregation algorithm
- Computable approximations for continuous-time Markov decision processes on Borel spaces based on empirical measures
- NDP methods for multi-chain MDPs
- New approximate dynamic programming algorithms for large-scale undiscounted Markov decision processes and their application to optimize a production and distribution system
- Simulation optimization algorithms for SMDPs with parameterized randomized stationary policies
- Strategic capacity decision-making in a stochastic manufacturing environment using real-time approximate dynamic programming
- What you should know about approximate dynamic programming
- Adaptive aggregation for reinforcement learning in average reward Markov decision processes
- The optimal control of just-in-time-based production and distribution systems and performance comparisons with optimized pull systems
- Simulation-based optimization of Markov reward processes
- Computable approximations for average Markov decision processes in continuous time
- A Sarsa() algorithm based on double-layer fuzzy reasoning
- Risk-Sensitive Reinforcement Learning via Policy Gradient Search
- Variance-penalized Markov decision processes: dynamic programming and reinforcement learning techniques
- An Adaptive Sampling Algorithm for Solving Markov Decision Processes
- Sampled fictitious play for approximate dynamic programming
- Simulation-based optimization of Markov decision processes: an empirical process theory approach
- A variable neighborhood search based algorithm for finite-horizon Markov decision processes
- Approximation of Markov decision processes with general state space
- A semi-Lagrangian approach for time and energy path planning optimization in static flow fields
- Multi-policy iteration with a distributed voting.
- Sleeping experts and bandits approach to constrained Markov decision processes
- A survey of some simulation-based algorithms for Markov decision processes
This page was built for publication: Simulation-based algorithms for Markov decision processes.
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q870662)