Optimal Scheduling of Entropy Regularizer for Continuous-Time Linear-Quadratic Reinforcement Learning
From MaRDI portal
Abstract: This work uses the entropy-regularised relaxed stochastic control perspective as a principled framework for designing reinforcement learning (RL) algorithms. Herein agent interacts with the environment by generating noisy controls distributed according to the optimal relaxed policy. The noisy policies, on the one hand, explore the space and hence facilitate learning but, on the other hand, introduce bias by assigning a positive probability to non-optimal actions. This exploration-exploitation trade-off is determined by the strength of entropy regularisation. We study algorithms resulting from two entropy regularisation formulations: the exploratory control approach, where entropy is added to the cost objective, and the proximal policy update approach, where entropy penalises the divergence of policies between two consecutive episodes. We analyse the finite horizon continuous-time linear-quadratic (LQ) RL problem for which both algorithms yield a Gaussian relaxed policy. We quantify the precise difference between the value functions of a Gaussian policy and its noisy evaluation and show that the execution noise must be independent across time. By tuning the frequency of sampling from relaxed policies and the parameter governing the strength of entropy regularisation, we prove that the regret, for both learning algorithms, is of the order (up to a logarithmic factor) over episodes, matching the best known result from the literature.
Cites work
- A modified MSA for stochastic control problems
- Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
- Adaptive continuous-time linear quadratic Gaussian control
- Compactification methods in the control of degenerate diffusions: existence of an optimal control
- Continuous‐time mean–variance portfolio selection: A reinforcement learning framework
- Entropy Regularization for Mean Field Games with Learning
- Exploratory LQG mean field games with entropy regularization
- High-dimensional probability. An introduction with applications in data science
- scientific article; zbMATH DE number 3567644 (Why is no real title available?)
- scientific article; zbMATH DE number 1325009 (Why is no real title available?)
- scientific article; zbMATH DE number 7307478 (Why is no real title available?)
- scientific article; zbMATH DE number 3283642 (Why is no real title available?)
- Mirror descent and nonlinear projected subgradient methods for convex optimization.
- On incomplete learning and certainty-equivalence control
- On the sample complexity of the linear quadratic regulator
- Regularity and stability of feedback relaxed controls
- Reinforcement learning and stochastic optimisation
- Reinforcement learning. An introduction
- The exact law of large numbers via Fubini extension and characterization of insurable risks
Cited in
(14)- Logarithmic regret bounds for continuous-time average-reward Markov decision processes
- Continuous-time optimal investment with portfolio constraints: a reinforcement learning approach
- Continuous time reinforcement learning: a random measure approach
- Fast policy learning for linear-quadratic control with entropy regularization
- Mirror descent for stochastic control problems with measure-valued controls
- Sublinear regret for a class of continuous-time linear-quadratic reinforcement learning problems
- Entropy annealing for policy mirror descent in continuous time and space
- Two system transformation data-driven algorithms for linear quadratic mean-field games
- Actor-critic learning algorithms for mean-field control with moment neural networks
- Continuous-time risk-sensitive reinforcement learning via quadratic variation penalty
- Accuracy of discretely sampled stochastic policies in continuous-time reinforcement learning
- -policy gradient for online pricing
- Logarithmic regret in the ergodic Avellaneda-Stoikov market making model
- Statistical learning with sublinear regret of propagator models
This page was built for publication: Optimal Scheduling of Entropy Regularizer for Continuous-Time Linear-Quadratic Reinforcement Learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6180253)