Unsynchronized decentralized Q-learning: two timescale analysis by persistence
From MaRDI portal
decentralized systemsindependent learnerslearning in gamesmultiagent reinforcement learningstochastic games
Applications of Markov chains and discrete-time Markov processes on general state spaces (social mobility, learning theory, industrial processes, etc.) (60J20) Stochastic games, stochastic differential games (91A15) Rationality and learning in game theory (91A26) Decentralized systems (93A14) Multi-agent systems (93A16)
Cites work
- 10.1162/1532443041827880
- Actor-Critic--Type Learning Algorithms for Markov Decision Processes
- Asynchronous stochastic approximation and Q-learning
- Convergence results for single-step on-policy reinforcement-learning algorithms
- Convergent multiple-timescales reinforcement learning algorithms in normal form games
- Cycles in adversarial regularized learning
- Decentralized Learning for Optimality in Stochastic Dynamic Teams and Games With Local Control and Global State Information
- Decentralized Q-Learning for Stochastic Teams and Games
- Discounted stochastic games with no stationary Nash equilibrium: two examples
- Error bounds for constant step-size \(Q\)-learning
- Global Nash convergence of Foster and Young's regret testing
- scientific article; zbMATH DE number 3205836 (Why is no real title available?)
- Independent learning in stochastic games
- Individual Q-Learning in Normal Form Games
- Multiagent learning using a variable learning rate
- On the nonconvergence of fictitious play in coordination games
- Payoff-based dynamics for multiplayer weakly acyclic games
- REINFORCEMENT LEARNING IN MARKOVIAN EVOLUTIONARY GAMES
- Stochastic imitation in finite games
This page was built for publication: Unsynchronized decentralized Q-learning: two timescale analysis by persistence
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6962174)