Dynamic shielding for reinforcement learning in black-box environments
From MaRDI portal
Computational learning theory (68Q32) Formal languages and automata (68Q45) Specification and verification (program logics, model checking, etc.) (68Q60) Learning and adaptive systems in artificial intelligence (68T05) Markov and semi-Markov decision processes (90C40) Control/observation systems governed by functional relations other than differential equations (such as hybrid and switching systems) (93C30)
Abstract: It is challenging to use reinforcement learning (RL) in cyber-physical systems due to the lack of safety guarantees during learning. Although there have been various proposals to reduce undesired behaviors during learning, most of these techniques require prior system knowledge, and their applicability is limited. This paper aims to reduce undesired behaviors during learning without requiring any prior system knowledge. We propose dynamic shielding: an extension of a model-based safe RL technique called shielding using automata learning. The dynamic shielding technique constructs an approximate system model in parallel with RL using a variant of the RPNI algorithm and suppresses undesired explorations due to the shield constructed from the learned model. Through this combination, potentially unsafe actions can be foreseen before the agent experiences them. Experiments show that our dynamic shield significantly decreases the number of undesired events during training.
Recommendations
- scientific article; zbMATH DE number 7559459
- Sim-to-lab-to-real: safe reinforcement learning with shielding and generalization guarantees
- Risk-aware shielding of partially observable Monte Carlo planning policies
- Safe reinforcement learning for dynamical games
- MULTIAGENT LEARNING FOR BLACK BOX SYSTEM REWARD FUNCTIONS
- Conditionally Elicitable Dynamic Risk Measures for Deep Reinforcement Learning
- The dynamics of generalized reinforcement learning
- The black box as a control for payoff-based learning in economic games
- Probabilistic guarantees for safe deep reinforcement learning
- Deep reinforcement learning with guaranteed performance. A Lyapunov-based approach
Cites work
- A comprehensive survey on safe reinforcement learning
- Deep reinforcement learning with temporal logics
- scientific article; zbMATH DE number 7626783 (Why is no real title available?)
- scientific article; zbMATH DE number 7559459 (Why is no real title available?)
- On the Construction of Fine Automata for Safety Properties
- On the Inference of Finite State Automata from Positive and Negative Data
- Run-time optimization for learned controllers through quantitative games
- Shield synthesis: runtime enforcement for reactive systems
This page was built for publication: Dynamic shielding for reinforcement learning in black-box environments
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6103158)