Safety-aware apprenticeship learning
From MaRDI portal
Abstract: Apprenticeship learning (AL) is a kind of Learning from Demonstration techniques where the reward function of a Markov Decision Process (MDP) is unknown to the learning agent and the agent has to derive a good policy by observing an expert's demonstrations. In this paper, we study the problem of how to make AL algorithms inherently safe while still meeting its learning objective. We consider a setting where the unknown reward function is assumed to be a linear combination of a set of state features, and the safety property is specified in Probabilistic Computation Tree Logic (PCTL). By embedding probabilistic model checking inside AL, we propose a novel counterexample-guided approach that can ensure safety while retaining performance of the learnt policy. We demonstrate the effectiveness of our approach on several challenging AL scenarios where safety is essential.
Recommendations
- A comprehensive survey on safe reinforcement learning
- Maximum causal entropy specification inference from demonstrations
- Safe exploration in model-based reinforcement learning using control barrier functions
- Supervisor synthesis of POMDP via automata learning
- A predictive safety filter for learning-based control of constrained nonlinear dynamical systems
Cited in
(9)- Counterexample-guided inductive synthesis for probabilistic systems
- Maximum causal entropy specification inference from demonstrations
- Interactive policy learning through confidence-based autonomy
- scientific article; zbMATH DE number 7559459 (Why is no real title available?)
- Apprenticeship learning in cognitive jamming
- Counterexample-driven synthesis for probabilistic program sketches
- Probabilistic counterexample guidance for safer reinforcement learning
- Policy-based primal-dual methods for concave CMDP with variance reduction
- Safe controller synthesis for nonlinear systems using Bayesian optimization enhanced reinforcement learning
This page was built for publication: Safety-aware apprenticeship learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6041137)