Derivative-free methods for policy optimization: guarantees for linear quadratic systems
From MaRDI portal
Abstract: We study derivative-free methods for policy optimization over the class of linear policies. We focus on characterizing the convergence rate of these methods when applied to linear-quadratic systems, and study various settings of driving noise and reward feedback. We show that these methods provably converge to within any pre-specified tolerance of the optimal policy with a number of zero-order evaluations that is an explicit polynomial of the error tolerance, dimension, and curvature properties of the problem. Our analysis reveals some interesting differences between the settings of additive driving noise and random initialization, as well as the settings of one-point and two-point reward feedback. Our theory is corroborated by extensive simulations of derivative-free methods on these systems. Along the way, we derive convergence rates for stochastic zero-order optimization algorithms when applied to a certain class of non-convex problems.
Recommendations
Cites work
- A Bound on Tail Probabilities for Quadratic Forms in Independent Random Variables
- A tail inequality for quadratic forms of subgaussian random vectors
- An Optimal Algorithm for Bandit and Zero-Order Convex Optimization with Two-Point Feedback
- Dynamic programming and optimal control. Vol. 1.
- End-to-end training of deep visuomotor policies
- Gradient methods for solving equations and inequalities
- scientific article; zbMATH DE number 3181381 (Why is no real title available?)
- scientific article; zbMATH DE number 1016965 (Why is no real title available?)
- scientific article; zbMATH DE number 3371284 (Why is no real title available?)
- Introduction to Stochastic Search and Optimization
- Linear Thompson sampling revisited
- On the sample complexity of the linear quadratic regulator
- Online convex optimization in the bandit setting: gradient descent without a gradient
- Optimal Rates for Zero-Order Convex Optimization: The Power of Two Function Evaluations
- Optimality of Fast-Matching Algorithms for Random Networks With Applications to Structural Controllability
- Optimization of Smooth Functions With Noisy Observations: Local Minimax Rates
- Probability. Theory and examples.
- Stochastic First- and Zeroth-Order Methods for Nonconvex Stochastic Programming
Cited in
(17)- Model-free linear quadratic regulator
- Controlled interacting particle algorithms for simulation-based reinforcement learning
- Model-free design of stochastic LQR controller from a primal-dual optimization perspective
- scientific article; zbMATH DE number 7625189 (Why is no real title available?)
- Tracking and Regret Bounds for Online Zeroth-Order Euclidean and Riemannian Optimization
- Policy Gradient Methods for the Noisy Linear Quadratic Regulator over a Finite Horizon
- Policy optimization for \(\mathcal{H}_2\) linear control with \(\mathcal{H}_\infty\) robustness guarantee: implicit regularization and global convergence
- Analysis of the optimization landscape of Linear Quadratic Gaussian (LQG) control
- Recent Theoretical Advances in Non-Convex Optimization
- Small errors in random zeroth-order optimization are imaginary
- Learning decentralized linear quadratic regulators with \(\sqrt{T}\) regret
- Sample complexity of the linear quadratic regulator: a reinforcement learning lens
- Policy gradient converges to the globally optimal policy for nearly linear-quadratic regulators
- A model-free first-order method for linear quadratic regulator with \(\tilde{O}(1/\varepsilon)\) sampling complexity
- Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
- On model-free learning over rate-limited channels: the linear quadratic regulator case
- Provable policy optimization for attitude control systems with unknown inertia
This page was built for publication: Derivative-free methods for policy optimization: guarantees for linear quadratic systems
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4969058)