Reinforcement learning with function approximation: from linear to nonlinear
The subject of this interesting paper is reinforcement learning with function approximation. The subject of reinforcement learning deals with how an agent can learn through interaction with the environment via an optimal policy that maximizes the long-term reward of the agent. When the problem contains a large number of states (often smooth and high-dimensional), a function approximation must be introduced to approximate the involved value or policy functions, and despite their success in many practical applications, for example video, computer vision and many others, a theoretical understanding of reinforcement learning algorithms with function approximation remains relatively limited, particularly when compared to the theoretical results in the tabular setting for the situation of a small number of states, which is much better understood theoretically.\N\NThe paper under review, in particular, reviews recent results on error analysis for reinforcement learning algorithms in linear or nonlinear approximation settings, emphasizing approximation error and estimation error/sample complexity, discussing in particular properties related to approximation error and presenting concrete conditions on reward and transition probability under which these properties hold.\N\NThe paper is well written with a good set of references.
- 10.1162/1532443041827907
- A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
- A priori estimates of the population risk for two-layer neural networks
- An L^2 analysis of reinforcement learning in high dimensions with kernel and neural network approximation
- An introduction to the theory of reproducing kernel Hilbert spaces
- Kolmogorov width decay and poor approximators in machine learning: shallow neural networks, random feature models and neural tangent kernels
- Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
- Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
- Multivariate \(L_{\infty}\) approximation in the worst case setting over reproducing kernel Hilbert spaces
- On the equivalence between kernel quadrature rules and random feature expansions
- On the theory of policy gradient methods: optimality, approximation, and distribution shift
- Optimal rates for the regularized least-squares algorithm
- Overcoming the curse of dimensionality in the numerical approximation of parabolic partial differential equations with gradient-dependent nonlinearities
- Perturbational complexity by distribution mismatch: a systematic analysis of reinforcement learning in reproducing kernel Hilbert space
- Regularized policy iteration with nonparametric function spaces
- Reinforcement learning. An introduction
- The Barron space and the flow-induced function spaces for neural network models
- Theory of Reproducing Kernels
- Towards a mathematical understanding of neural network-based machine learning: what we know and what we don't
This page was built for publication: Reinforcement learning with function approximation: from linear to nonlinear
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6955719)