Performance gradient estimation for the very large finite Markov chains
From MaRDI portal
Applications of Markov chains and discrete-time Markov processes on general state spaces (social mobility, learning theory, industrial processes, etc.) (60J20) Applications of Markov renewal processes (reliability, queueing networks, etc.) (60K20) Queueing theory (aspects of probability theory) (60K25)
Recommendations
Cited in
(10)- Gradient estimates for the performance of Markov chains and discrete event processes
- On performance potentials and conditional Monte Carlo for gradient estimation for Markov chains
- A time aggregation approach to Markov decision processes
- Optimization via simulation: A review
- A basic formula for performance gradient estimation of semi-Markov decision processes
- A unified approach to time-aggregated Markov decision processes
- The control of a two-level Markov decision process by time aggregation
- Solving average cost Markov decision processes by means of a two-phase time aggregation algorithm
- Likelihood ratio gradient estimation for steady-state parameters
- On-line policy gradient estimation with multi-step sampling
This page was built for publication: Performance gradient estimation for the very large finite Markov chains
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3986172)