Faster algorithm and sharper analysis for constrained Markov decision process
From MaRDI portal
Cites work
- A comprehensive survey on safe reinforcement learning
- A General Safety Framework for Learning-Based Control in Uncertain Robotic Systems
- First-order and stochastic optimization methods for machine learning
- High-dimensional statistics. A non-asymptotic viewpoint
- scientific article; zbMATH DE number 1348599 (Why is no real title available?)
- Information-based complexity of linear operator equations
- Lower complexity bounds of first-order methods for convex-concave bilinear saddle-point problems
- Markov chains and mixing times. With a chapter on ``Coupling from the past by James G. Propp and David B. Wilson.
- Policy mirror descent for reinforcement learning: linear convergence, new sampling complexity, and generalized problem classes
- Reinforcement learning. An introduction
- Risk-constrained reinforcement learning with percentile risk criteria
- Sensitivity and convergence of uniformly ergodic Markov chains
Cited in
(4)- A primal-dual policy iteration algorithm for constrained Markov decision processes
- Stochastic optimization under hidden convexity
- Policy-based primal-dual methods for concave CMDP with variance reduction
- Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
This page was built for publication: Faster algorithm and sharper analysis for constrained Markov decision process
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6988660)