Estimation and approximation bounds for gradient-based reinforcement learning
From MaRDI portal
Recommendations
- Approximate gradient methods in policy-space optimization of Markov reward processes
- Restricted gradient-descent algorithm for value-function approximation in reinforcement learning
- Reinforcement learning with approximation spaces
- Reinforcement learning via approximation of the Q-function
- Analysis and improvement of policy gradient estimation
- On the theory of policy gradient methods: optimality, approximation, and distribution shift
- Near-optimal regret bounds for reinforcement learning
- Expected policy gradients for reinforcement learning
- On tight bounds for function approximation error in risk-sensitive reinforcement learning
Cites work
- scientific article; zbMATH DE number 1753152 (Why is no real title available?)
- scientific article; zbMATH DE number 3443893 (Why is no real title available?)
- Learning dynamical systems in a stationary environment
- Minimum complexity regression estimation with weakly dependent observations
- Neural Network Learning
- Nonparametric time series prediction through adaptive model selection
- OnActor-Critic Algorithms
- Probability Inequalities for Sums of Bounded Random Variables
- Sensitivity Analysis for Simulations via Likelihood Ratios
- Simple statistical gradient-following algorithms for connectionist reinforcement learning
Cited in
(8)- Relative loss bounds for temporal-difference learning
- Exploiting random walks for learning
- Attainability of boundary points under reinforcement learning
- scientific article; zbMATH DE number 1753152 (Why is no real title available?)
- scientific article; zbMATH DE number 1753153 (Why is no real title available?)
- On Overfitting and Asymptotic Bias in Batch Reinforcement Learning with Partial Observability
- On the sample complexity of actor-critic method for reinforcement learning with function approximation
- Finite-time analysis of natural actor-critic for POMDPs
This page was built for publication: Estimation and approximation bounds for gradient-based reinforcement learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q1604222)