Dual control for approximate Bayesian reinforcement learning
From MaRDI portal
Abstract: Control of non-episodic, finite-horizon dynamical systems with uncertain dynamics poses a tough and elementary case of the exploration-exploitation trade-off. Bayesian reinforcement learning, reasoning about the effect of actions and future observations, offers a principled solution, but is intractable. We review, then extend an old approximate approach from control theory---where the problem is known as dual control---in the context of modern regression methods, specifically generalized linear regression. Experiments on simulated systems show that this framework offers a useful approximation to the intractable aspects of Bayesian RL, producing structured exploration strategies that differ from standard RL approaches. We provide simple examples for the use of this framework in (approximate) Gaussian process regression and feedforward neural networks for the control of exploration.
Recommendations
- Bayesian exploration for approximate dynamic programming
- Bayesian Reinforcement Learning with Exploration
- A model for system uncertainty in reinforcement learning
- A Bayesian approach for learning and planning in partially observable Markov decision processes
- Bayesian optimistic Kullback-Leibler exploration
Cited in
(13)- A model for system uncertainty in reinforcement learning
- Dual control of linearly parameterised models via prediction of posterior densities
- Reliable control based on dual control for ARMAX system with abrupt faults
- Dual control for exploitation and exploration (DCEE) in autonomous search
- Fully probabilistic design of strategies with estimator
- An active exploration method for data efficient reinforcement learning
- Controller exploitation-exploration reinforcement learning architecture for computing near-optimal policies
- Bayesian optimistic Kullback-Leibler exploration
- Bayesian reinforcement learning: a survey
- Variance regularization in sequential Bayesian optimization
- Bayesian exploration for approximate dynamic programming
- Robust reinforcement learning with Bayesian optimisation and quadrature
- A Bayesian reinforcement learning approach in Markov games for computing near-optimal policies
This page was built for publication: Dual control for approximate Bayesian reinforcement learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2834442)