Model-free design of stochastic LQR controller from a primal-dual optimization perspective
From MaRDI portal
Abstract: To further understand the underlying mechanism of various reinforcement learning (RL) algorithms and also to better use the optimization theory to make further progress in RL, many researchers begin to revisit the linear-quadratic regulator (LQR) problem, whose setting is simple and yet captures the characteristics of RL. Inspired by this, this work is concerned with the model-free design of stochastic LQR controller for linear systems subject to Gaussian noises, from the perspective of both RL and primal-dual optimization. From the RL perspective, we first develop a new model-free off-policy policy iteration (MF-OPPI) algorithm, in which the sampled data is repeatedly used for updating the policy to alleviate the data-hungry problem to some extent. We then provide a rigorous analysis for algorithm convergence by showing that the involved iterations are equivalent to the iterations in the classical policy iteration (PI) algorithm. From the perspective of optimization, we first reformulate the stochastic LQR problem at hand as a constrained non-convex optimization problem, which is shown to have strong duality. Then, to solve this non-convex optimization problem, we propose a model-based primal-dual (MB-PD) algorithm based on the properties of the resulting Karush-Kuhn-Tucker (KKT) conditions. We also give a model-free implementation for the MB-PD algorithm by solving a transformed dual feasibility condition. More importantly, we show that the dual and primal update steps in the MB-PD algorithm can be interpreted as the policy evaluation and policy improvement steps in the PI algorithm, respectively. Finally, we provide one simulation example to show the performance of the proposed algorithms.
Recommendations
- Model-free LQR design by Q-function learning
- Model-free linear quadratic regulator
- A reinforcement learning-based scheme for direct adaptive optimal control of linear stochastic systems
- Robust reinforcement learning for stochastic linear quadratic control with multiplicative noise
- Convergence results for an averaged LQR problem with applications to reinforcement learning
Cites work
- \(\mathrm{H}_\infty\) control of linear discrete-time systems: off-policy reinforcement learning
- Computational adaptive optimal control for continuous-time linear systems with completely unknown dynamics
- Deep reinforcement learning for wireless sensor scheduling in cyber-physical systems
- Derivative-free methods for policy optimization: guarantees for linear quadratic systems
- Discrete-time linear systems. Theory and design with applications.
- scientific article; zbMATH DE number 2107836 (Why is no real title available?)
- On the sample complexity of the linear quadratic regulator
- On the theory of policy gradient methods: optimality, approximation, and distribution shift
- Optimal adaptive control and differential games by reinforcement learning principles
- Primal-Dual Q-Learning Framework for LQR Design
- Stability Analysis of Discrete-Time Infinite-Horizon Optimal Control With Discounted Cost
- Stochastic linear-quadratic control via semidefinite programming
Cited in
(12)- Free finite horizon LQR: a bilevel perspective and its application to model predictive control
- Model-free optimal control of discrete-time systems with additive and multiplicative noises
- Primal-Dual Q-Learning Framework for LQR Design
- An inertial neural network approach for loco-manipulation trajectory tracking of mobile robot with redundant manipulator
- An adaptive dynamic programming-based algorithm for infinite-horizon linear quadratic stochastic optimal control problems
- Lossless convexification and duality
- A data-ensemble-based approach for sample-efficient LQ control of linear time-varying systems
- Model-free stochastic linear quadratic design by semidefinite programming
- Model-free approximate dynamic programming for stochastic zero-sum games: algorithm design and analysis
- Stochastic linear quadratic optimal control for continuous-time systems via reinforcement learning
- Direct data-driven discounted infinite horizon linear quadratic regulator with robustness guarantees
- Stochastic H₂/H_ off-policy reinforcement learning tracking control for linear discrete-time systems with multiplicative noises
This page was built for publication: Model-free design of stochastic LQR controller from a primal-dual optimization perspective
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2125546)