Multiagent Online Learning in Time-Varying Games
From MaRDI portal
Abstract: We examine the long-run behavior of multi-agent online learning in games that evolve over time. Specifically, we focus on a wide class of policies based on mirror descent, and we show that the induced sequence of play (a) converges to Nash equilibrium in time-varying games that stabilize in the long run to a strictly monotone limit; and (b) it stays asymptotically close to the evolving equilibrium of the sequence of stage games (assuming they are strongly monotone). Our results apply to both gradient-based and payoff-based feedback - i.e., the "bandit feedback" case where players only get to observe the payoffs of their chosen actions.
Recommendations
- Learning in games with continuous action sets and unknown payoff functions
- A unified stochastic approximation framework for learning in games
- Fast convergence of optimistic gradient ascent in network zero-sum extensive form games
- On the robustness of learning in games with stochastically perturbed payoff observations
- Learning in nonatomic games. I: Finite action spaces and population games
Cited in
(4)- Nested replicator dynamics, nested logit choice, and similarity-based learning
- Derivative-free stochastic bilevel optimization for inverse problems
- Distributed Nash equilibrium seeking with a dynamic set of players
- A distributed iterative Tikhonov method for networked monotone stochastic and hierarchical aggregative games
This page was built for publication: Multiagent Online Learning in Time-Varying Games
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6199277)