Global optimality guarantees for policy gradient methods
From MaRDI portal
Recommendations
- On the theory of policy gradient methods: optimality, approximation, and distribution shift
- Global convergence of policy gradient methods to (almost) locally optimal policies
- Policy gradient methods for discrete time linear quadratic regulator with random parameters
- Fast global convergence of natural policy gradient methods with entropy regularization
Cited in
(12)- On the theory of policy gradient methods: optimality, approximation, and distribution shift
- Occupancy information ratio: infinite-horizon, information-directed, parameterized policy search
- Hierarchical dynamic graphical games for optimal leader-follower consensus control
- Stochastic approximation and reinforcement learning: the interface and a little beyond
- Fast policy learning for linear-quadratic control with entropy regularization
- On the convergence of projected policy gradient for any constant step sizes
- Search or split: policy gradient with adaptive policy space
- Convergence of natural policy gradient for a family of infinite-state queueing MDPs
- Fisher-Rao gradient flows of linear programs and state-action natural policy gradients
- Learning collusive strategies with function approximation algorithms
- Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
- Policy optimization over general state and action spaces
This page was built for publication: Global optimality guarantees for policy gradient methods
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6655175)