Policy iteration based on stochastic factorization
From MaRDI portal
Recommendations
- Factored value iteration converges
- Stochastic dynamic programming with factored representations
- Generalized polynomial approximations in Markovian decision processes
- State reduction in a Markov decision process
- The simplex and policy-iteration methods are strongly polynomial for the Markov decision problem with a fixed discount rate
Cited in
(8)- Truncated policy iteration methods
- An incremental off-policy search in a model-free Markov decision process using a single sample path
- A numerical study of Markov decision process algorithms for multi-component replacement problems
- Factored value iteration converges
- 10.1162/1532443041827907
- scientific article; zbMATH DE number 2247690 (Why is no real title available?)
- Offline reinforcement learning in large state spaces: algorithms and guarantees
- The factored policy-gradient planner
This page was built for publication: Policy iteration based on stochastic factorization
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2878742)