Extreme state aggregation beyond MDPs
From MaRDI portal
Abstract: We consider a Reinforcement Learning setup where an agent interacts with an environment in observation-reward-action cycles without any (esp. MDP) assumptions on the environment. State aggregation and more generally feature reinforcement learning is concerned with mapping histories/raw-states to reduced/aggregated states. The idea behind both is that the resulting reduced process (approximately) forms a small stationary finite-state MDP, which can then be efficiently solved or learnt. We considerably generalize existing aggregation results by showing that even if the reduced process is not an MDP, the (q-)value functions and (optimal) policies of an associated MDP with same state-space size solve the original problem, as long as the solution can approximately be represented as a function of the reduced states. This implies an upper bound on the required state space size that holds uniformly for all RL problems. It may also explain why RL algorithms designed for MDPs sometimes perform well beyond MDPs.
Recommendations
- Extreme state aggregation beyond Markov decision processes
- Selecting near-optimal approximate state representations in reinforcement learning
- Adaptive aggregation for reinforcement learning in average reward Markov decision processes
- Relative value iteration algorithm with soft state aggregation
- Unsupervised basis function adaptation for reinforcement learning
Cited in
(6)- Structure in the space of value functions
- Selecting near-optimal approximate state representations in reinforcement learning
- Chasing Ghosts: Competing with Stateful Policies
- Extreme state aggregation beyond Markov decision processes
- The Benefits of State Aggregation with Extreme-Point Weighting for Assemble-to-Order Systems
- On State Aggregation to Approximate Complex Value Functions in Large-Scale Markov Decision Processes
This page was built for publication: Extreme state aggregation beyond MDPs
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2938732)