Reward machines: exploiting reward function structure in reinforcement learning
From MaRDI portal
Recommendations
- Faithful and Effective Reward Schemes for Model-Free Reinforcement Learning of Omega-Regular Objectives
- Induction and exploitation of subgoal automata for reinforcement learning
- Reinforcement learning in sparse-reward environments with hindsight policy gradients
- Model-based average reward reinforcement learning
- Reinforcement learning of non-Markov decision processes
Cites work
- \({\mathcal Q}\)-learning
- Approximate Value Iteration with Temporally Extended Actions
- Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
- Deep reinforcement learning with temporal logics
- scientific article; zbMATH DE number 3664335 (Why is no real title available?)
- scientific article; zbMATH DE number 1560499 (Why is no real title available?)
- Induction and exploitation of subgoal automata for reinforcement learning
- The simplex method is strongly polynomial for deterministic Markov decision processes
- Transfer of learning by composing solutions of elemental sequential tasks
Cited in
(15)- Detect, understand, act: a neuro-symbolic hierarchical reinforcement learning framework
- Induction and exploitation of subgoal automata for reinforcement learning
- Learning reward machines: a study in partially observable reinforcement learning
- An impossibility result in automata-theoretic reinforcement learning
- A framework for transforming specifications in reinforcement learning
- Reward tampering problems and solutions in reinforcement learning: a causal influence diagram perspective
- Faithful and Effective Reward Schemes for Model-Free Reinforcement Learning of Omega-Regular Objectives
- A neurosymbolic cognitive architecture framework for handling novelties in open worlds
- Regular decision processes
- Joint learning of reward machines and policies in environments with partially known semantics
- Shaping reward signals in reinforcement learning using constraint programming
- Designing equilibria in concurrent games with social welfare and temporal logic constraints
- Synthesising reward machines for cooperative multi-agent reinforcement learning
- Average reward reinforcement learning for omega-regular and mean-payoff objectives
- Framing reinforcement learning from human reward: reward positivity, temporal discounting, episodicity, and performance
This page was built for publication: Reward machines: exploiting reward function structure in reinforcement learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5026256)