Evaluation and learning in two-player symmetric games via best and better responses
From MaRDI portal
Publication:6095620
Abstract: Artificial intelligence and robotic competitions are accompanied by a class of game paradigms in which each player privately commits a strategy to a game system which simulates the game using the collected joint strategy and then returns payoffs to players. This paper considers the strategy commitment for two-player symmetric games in which the players' strategy spaces are identical and their payoffs are symmetric. First, we introduce two digraph-based metrics at a meta-level for strategy evaluation in two-agent reinforcement learning, grounded on sink equilibrium. The metrics rank the strategies of a single player and determine the set of strategies which are preferred for the private commitment. Then, in order to find the preferred strategies under the metrics, we propose two variants of the classical learning algorithm self-play, called strictly best-response and weakly better-response self-plays. By modeling learning processes as walks over joint-strategy response digraphs, we prove that the learnt strategies by two variants are preferred under two metrics, respectively. The preferred strategies under both two metrics are identified and adjacency matrices induced by one metric and one variant are connected. Finally, simulations are provided to illustrate the results.
Recommendations
- AWESOME: a general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents
- Memory-two strategies forming symmetric mutual reinforcement learning equilibrium in repeated prisoners' dilemma game
- Symmetric equilibrium of multi-agent reinforcement learning in repeated prisoner's dilemma
- Effective choice in all the symmetric \(2\times 2\) games
- Two-speed evolution of strategies and preferences in symmetric games
Cites work
- 10.1162/1532443041827880
- Control of Learning in Anticoordination Network Games
- Convergent learning algorithms for unknown reward games
- Decentralized Learning for Optimality in Stochastic Dynamic Teams and Games With Local Control and Global State Information
- Decentralized Q-Learning for Stochastic Teams and Games
- Distributed Nash Equilibrium Seeking by a Consensus Based Approach
- Dynamic NE Seeking for Multi-Integrator Networked Agents With Disturbance Rejection
- Fast Convergence in Semianonymous Potential Games
- scientific article; zbMATH DE number 5301288 (Why is no real title available?)
- scientific article; zbMATH DE number 3204219 (Why is no real title available?)
- scientific article; zbMATH DE number 3069635 (Why is no real title available?)
- Nash equilibrium seeking for N-coalition noncooperative games
- Policy Evaluation and Seeking for Multiagent Reinforcement Learning via Best Response
- Potential field hierarchical reinforcement learning approach for target search by multi-AUV in 3-D underwater environments
- Potential games
- Proximal Dynamics in Multiagent Network Games
- Revisiting log-linear learning: asynchrony, completeness and payoff-based implementation
- The Impact of Complex and Informed Adversarial Behavior in Graphical Coordination Games
This page was built for publication: Evaluation and learning in two-player symmetric games via best and better responses
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6095620)