Safe Exploration of State and Action Spaces in Reinforcement Learning
From MaRDI portal
Abstract: In this paper, we consider the important problem of safe exploration in reinforcement learning. While reinforcement learning is well-suited to domains with complex transition dynamics and high-dimensional state-action spaces, an additional challenge is posed by the need for safe and efficient exploration. Traditional exploration techniques are not particularly useful for solving dangerous tasks, where the trial and error process may lead to the selection of actions whose execution in some states may result in damage to the learning system (or any other system). Consequently, when an agent begins an interaction with a dangerous and high-dimensional state-action space, an important question arises; namely, that of how to avoid (or at least minimize) damage caused by the exploration of the state-action space. We introduce the PI-SRL algorithm which safely improves suboptimal albeit robust behaviors for continuous state and action control tasks and which efficiently learns from the experience gained from the environment. We evaluate the proposed method in four complex tasks: automatic car parking, pole-balancing, helicopter hovering, and business management.
Recommendations
- Safe reinforcement learning for continuous spaces through Lyapunov-constrained behavior
- Safe exploration in model-based reinforcement learning using control barrier functions
- Safe reinforcement learning for dynamical games
- Planning for potential: efficient safe reinforcement learning
- Safe Reinforcement Learning Using Robust MPC
- A comprehensive survey on safe reinforcement learning
- Probabilistic guarantees for safe deep reinforcement learning
- Safety-constrained reinforcement learning with a distributional safety critic
- scientific article; zbMATH DE number 7559459
Cited in
(22)- k-Certainty Exploration Method: an action selector to identify the environment in reinforcement learning
- Reinforcement learning endowed with safe veto policies to learn the control of linked-multicomponent robotic systems
- Risk-averse autonomous systems: a brief history and recent developments from the perspective of optimal control
- Safe exploration in model-based reinforcement learning using control barrier functions
- Safe global optimization of expensive noisy black-box functions in the -Lipschitz framework
- Planning for potential: efficient safe reinforcement learning
- An active exploration method for data efficient reinforcement learning
- Safe reinforcement learning for continuous spaces through Lyapunov-constrained behavior
- Probabilistic inference for determining options in reinforcement learning
- Robust reinforcement learning with Bayesian optimisation and quadrature
- Deep exploration via randomized value functions
- Blind spot detection for safe sim-to-real transfer
- A comprehensive survey on safe reinforcement learning
- Bayesian optimization with safety constraints: safe and automatic parameter tuning in robotics
- Risk-aware controller for autonomous vehicles using model-based collision prediction and reinforcement learning
- Explicit explore, exploit, or escape \((E^4)\): near-optimal safety-constrained reinforcement learning in polynomial time
- A novel policy based on action confidence limit to improve exploration efficiency in reinforcement learning
- Certified reinforcement learning with logic guidance
- Discovering diverse solutions in deep reinforcement learning by maximizing state-action-based mutual information
- Probabilistic counterexample guidance for safer reinforcement learning
- All-time safety and sample-efficient meta update for online safe meta reinforcement learning under Markov task transition
- Safe navigation in adversarial environments
This page was built for publication: Safe Exploration of State and Action Spaces in Reinforcement Learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4899127)