Efficient sample reuse in policy gradients with parameter-based exploration
From MaRDI portal
(Redirected from Publication:5378202)
Abstract: The policy gradient approach is a flexible and powerful reinforcement learning method particularly for problems with continuous actions such as robot control. A common challenge in this scenario is how to reduce the variance of policy gradient estimates for reliable policy updates. In this paper, we combine the following three ideas and give a highly effective policy gradient method: (a) the policy gradients with parameter based exploration, which is a recently proposed policy search method with low variance of gradient estimates, (b) an importance sampling technique, which allows us to reuse previously gathered data in a consistent way, and (c) an optimal baseline, which minimizes the variance of gradient estimates with their unbiasedness being maintained. For the proposed method, we give theoretical analysis of the variance of gradient estimates and show its usefulness through extensive experiments.
Recommendations
- Analysis and improvement of policy gradient estimation
- Reward-weighted regression with sample reuse for direct policy search in reinforcement learning
- Expected policy gradients for reinforcement learning
- Importance sampling techniques for policy optimization
- Variance reduction techniques for gradient estimates in reinforcement learning
Cites work
- Analysis and improvement of policy gradient estimation
- Approximate dynamic programming with a fuzzy parameterization
- scientific article; zbMATH DE number 3906232 (Why is no real title available?)
- scientific article; zbMATH DE number 854710 (Why is no real title available?)
- Improving predictive inference under covariate shift by weighting the log-likelihood function
- Real-time reinforcement learning by sequential actor-critics and experience replay
- Reinforcement learning. An introduction
- Reward-weighted regression with sample reuse for direct policy search in reinforcement learning
- Simple statistical gradient-following algorithms for connectionist reinforcement learning
- Variance reduction techniques for gradient estimates in reinforcement learning
Cited in
(17)- Adaptive importance sampling for value function approximation in off-policy reinforcement learning
- Efficient exploration through active learning for value function approximation in reinforcement learning
- Importance sampling in reinforcement learning with an estimated behavior policy
- Model-based reinforcement learning with dimension reduction
- An active exploration method for data efficient reinforcement learning
- Reward-weighted regression with sample reuse for direct policy search in reinforcement learning
- A generalized path integral control approach to reinforcement learning
- Recurrent policy gradients
- Analysis and improvement of policy gradient estimation
- Expected policy gradients for reinforcement learning
- Importance sampling techniques for policy optimization
- Reinforcement learning in sparse-reward environments with hindsight policy gradients
- An Incremental Fast Policy Search Using a Single Sample Path
- Policy search for active fault diagnosis with partially observable state
- Deep ensemble reinforcement learning with multiple deep deterministic policy gradient algorithm
- Learning under nonstationarity: covariate shift and class-balance change
- Model-based policy gradients with parameter-based exploration by least-squares conditional density estimation
This page was built for publication: Efficient sample reuse in policy gradients with parameter-based exploration
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5378202)