Variance reduction techniques for gradient estimates in reinforcement learning
From MaRDI portal
Recommendations
- Geometric variance reduction in Markov chains: application to value function and gradient estimation
- Using Gaussian processes for variance reduction in policy gradient algorithms
- scientific article; zbMATH DE number 1753152
- Bayesian policy gradient and actor-critic algorithms
- scientific article; zbMATH DE number 1753153
Cited in
(25)- Natural actor-critic algorithms
- Importance sampling in reinforcement learning with an estimated behavior policy
- TD-regularized actor-critic methods
- Optimised graded metamaterials for mechanical energy confinement and amplification via reinforcement learning
- Using Gaussian processes for variance reduction in policy gradient algorithms
- Adaptive playouts for online learning of policies during Monte Carlo tree search
- Geometric variance reduction in Markov chains: application to value function and gradient estimation
- Learning to control a structured-prediction decoder for detection of HTTP-layer DDoS attackers
- Analysis and improvement of policy gradient estimation
- scientific article; zbMATH DE number 1753152 (Why is no real title available?)
- scientific article; zbMATH DE number 7014219 (Why is no real title available?)
- Expected policy gradients for reinforcement learning
- Monte Carlo gradient estimation in machine learning
- Reinforcement learning in sparse-reward environments with hindsight policy gradients
- Global convergence of policy gradient methods to (almost) locally optimal policies
- Deep Reinforcement Learning: A State-of-the-Art Walkthrough
- Efficient sample reuse in policy gradients with parameter-based exploration
- On-line policy gradient estimation with multi-step sampling
- Optimistic reinforcement learning by forward Kullback-Leibler divergence optimization
- Variational actor-critic algorithms,
- Personalized dynamic treatment regimes in continuous time: a Bayesian approach for optimizing clinical decisions with timing
- A Bayesian decision framework for optimizing sequential combination antiretroviral therapy in people with HIV
- Scalable Control Variates for Monte Carlo Methods Via Stochastic Optimization
- Leveraging randomized smoothing for optimal control of nonsmooth dynamical systems
- The factored policy-gradient planner
This page was built for publication: Variance reduction techniques for gradient estimates in reinforcement learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3093234)