Omega-Regular Objectives in Model-Free Reinforcement Learning
From MaRDI portal
Abstract: We provide the first solution for model-free reinforcement learning of {omega}-regular objectives for Markov decision processes (MDPs). We present a constructive reduction from the almost-sure satisfaction of {omega}-regular objectives to an almost- sure reachability problem and extend this technique to learning how to control an unknown model so that the chance of satisfying the objective is maximized. A key feature of our technique is the compilation of {omega}-regular properties into limit- deterministic Buechi automata instead of the traditional Rabin automata; this choice sidesteps difficulties that have marred previous proposals. Our approach allows us to apply model-free, off-the-shelf reinforcement learning algorithms to compute optimal strategies from the observations of the MDP. We present an experimental evaluation of our technique on benchmark learning problems.
Recommendations
- Faithful and Effective Reward Schemes for Model-Free Reinforcement Learning of Omega-Regular Objectives
- Model-Free Reinforcement Learning for Lexicographic Omega-Regular Objectives
- Learning-based mean-payoff optimization in an unknown MDP under omega-regular constraints
- Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning
- Near-optimal regret bounds for reinforcement learning
- Model-based reinforcement learning for approximate optimal regulation
- Choquet Regularization for Continuous-Time Reinforcement Learning
- What is decidable about partially observable Markov decision processes with omega-regular objectives
Cited in
(16)- Deep reinforcement learning with temporal logics
- Probabilistic guarantees for safe deep reinforcement learning
- Minimum Attention Controller Synthesis for Omega-Regular Objectives
- An impossibility result in automata-theoretic reinforcement learning
- Alternating good-for-MDPs automata
- Specification-guided reinforcement learning
- Faithful and Effective Reward Schemes for Model-Free Reinforcement Learning of Omega-Regular Objectives
- PAC Statistical Model Checking of Mean Payoff in Discrete- and Continuous-Time MDP
- Model-Free Reinforcement Learning for Lexicographic Omega-Regular Objectives
- Policy synthesis and reinforcement learning for discounted LTL
- Joint learning of reward machines and policies in environments with partially known semantics
- PAC statistical model checking of mean payoff in discrete- and continuous-time MDP
- Deciding what is good-for-MDPs
- Multi-objective -regular reinforcement learning
- Average reward reinforcement learning for omega-regular and mean-payoff objectives
- Enforcing almost-sure reachability in POMDPs
This page was built for publication: Omega-Regular Objectives in Model-Free Reinforcement Learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6091336)