Optimal and approximate Q-value functions for decentralized POMDPS
From MaRDI portal
Recommendations
Cited in
(14)- A unified framework for stochastic optimization
- Multi-agent reinforcement learning algorithm to solve a partially-observable multi-agent problem in disaster response
- Improving coordination in small-scale multi-agent deep reinforcement learning through memory-driven communication
- A leader-follower partially observed, multiobjective Markov game
- Optimally solving Dec-POMDPs as continuous-state MDPs
- An investigation into mathematical programming for finite horizon decentralized POMDPS
- The cross-entropy method for policy search in decentralized POMDPs
- Centralized Optimization for Dec-POMDPs Under the Expected Average Reward Criterion
- Controlling a Fleet of Unmanned Aerial Vehicles to Collect Uncertain Information in a Threat Environment
- Byzantine-Resilient Decentralized Policy Evaluation With Linear Function Approximation
- Monotonic value function factorisation for deep multi-agent reinforcement learning
- Online planning for multi-agent systems with bounded communication
- A sufficient statistic for influence in structured multiagent environments
- On centralized critics in multi-agent reinforcement learning
This page was built for publication: Optimal and approximate Q-value functions for decentralized POMDPS
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3624133)