Verifiably Safe Off-Model Reinforcement Learning
From MaRDI portal
Abstract: The desire to use reinforcement learning in safety-critical settings has inspired a recent interest in formal methods for learning algorithms. Existing formal methods for learning and optimization primarily consider the problem of constrained learning or constrained optimization. Given a single correct model and associated safety constraint, these approaches guarantee efficient learning while provably avoiding behaviors outside the safety constraint. Acting well given an accurate environmental model is an important pre-requisite for safe learning, but is ultimately insufficient for systems that operate in complex heterogeneous environments. This paper introduces verification-preserving model updates, the first approach toward obtaining formal safety guarantees for reinforcement learning in settings where multiple environmental models must be taken into account. Through a combination of design-time model updates and runtime model falsification, we provide a first approach toward obtaining formal safety proofs for autonomous systems acting in heterogeneous environments.
Recommendations
- Learning through imitation by using formal verification
- Safe exploration in model-based reinforcement learning using control barrier functions
- A comprehensive survey on safe reinforcement learning
- Off‐policy model‐based end‐to‐end safe reinforcement learning
- A learner-verifier framework for neural network controllers and certificates of stochastic systems
Cited in
(6)- Deep reinforcement learning with temporal logics
- Planning for potential: efficient safe reinforcement learning
- Learning through imitation by using formal verification
- Verification-guided programmatic controller synthesis
- Probabilistic reach-avoid for Bayesian neural networks
- Adaptive network approach to exploration-exploitation trade-off in reinforcement learning
This page was built for publication: Verifiably Safe Off-Model Reinforcement Learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6091337)