Learning in games via reinforcement and regularization
From MaRDI portal
Abstract: We investigate a class of reinforcement learning dynamics where players adjust their strategies based on their actions' cumulative payoffs over time - specifically, by playing mixed strategies that maximize their expected cumulative payoff minus a regularization term. A widely studied example is exponential reinforcement learning, a process induced by an entropic regularization term which leads mixed strategies to evolve according to the replicator dynamics. However, in contrast to the class of regularization functions used to define smooth best responses in models of stochastic fictitious play, the functions used in this paper need not be infinitely steep at the boundary of the simplex; in fact, dropping this requirement gives rise to an important dichotomy between steep and nonsteep cases. In this general framework, we extend several properties of exponential learning, including the elimination of dominated strategies, the asymptotic stability of strict Nash equilibria, and the convergence of time-averaged trajectories in zero-sum games with an interior Nash equilibrium.
Recommendations
Cites work
- ``Evolutionary selection dynamic in games: Convergence and limit properties
- A continuous-time approach to online optimization
- A note on best response dynamics.
- A payoff-based learning procedure and its application to traffic games
- A Simple Adaptive Procedure Leading to Correlated Equilibrium
- Adaptive dynamics and evolutionary stability
- Adaptive game playing using multiplicative weights
- Attainability of boundary points under reinforcement learning
- Barrier Operators and Associated Gradient-Like Dynamical Systems for Constrained Minimization Problems
- Domination or equilibrium
- Escort evolutionary game theory
- Evolutionary Games and Population Dynamics
- Evolutionary Games in Economics
- Evolutionary stability in asymmetric games
- Exponential weight algorithm in continuous time
- Free-Steering Relaxation Methods for Problems with Strictly Convex Costs and Linear Constraints
- Hessian Riemannian Gradient Flows in Convex Programming
- Higher order game dynamics
- scientific article; zbMATH DE number 4141836 (Why is no real title available?)
- scientific article; zbMATH DE number 918596 (Why is no real title available?)
- Individual Q-Learning in Normal Form Games
- Inertial game dynamics and applications to constrained optimization
- Learning through reinforcement and replicator dynamics
- Learning, matching, and aggregation
- Mirror descent and nonlinear projected subgradient methods for convex optimization.
- No-regret dynamics and fictitious play
- On the convergence of reinforcement learning
- On the Global Convergence of Stochastic Fictitious Play
- Online learning and online convex optimization
- Optimal properties of stimulus-response learning models.
- Penalty-regulated dynamics and robust learning procedures in games
- Perturbed variations of penalty function methods. Example: Projective SUMT
- Possible generalization of Boltzmann-Gibbs statistics.
- Primal-dual subgradient methods for convex problems
- Projected Dynamical Systems in the Formulation, Stability Analysis, and Computation of Fixed-Demand Traffic Network Equilibria
- Quantal response equilibria for normal form games
- Riemannian game dynamics
- Social Stability and Equilibrium
- Stochastic Approximations and Differential Inclusions
- The emergence of rational behavior in the presence of stochastic perturbations
- The Nonlinear Geometry of Linear Programming. I Affine and Projective Scaling Trajectories
- The projection dynamic and the geometry of population games
- The projection dynamic and the replicator dynamic
- The weighted majority algorithm
- Time Average Replicator and Best-Reply Dynamics
- Two Competing Models of How People Learn in Games
Cited in
(46)- Riemannian game dynamics
- Evolutionary game theory: a renaissance
- Learning in games with continuous action sets and unknown payoff functions
- On the convergence of reinforcement learning
- Population games and discrete optimal transport
- Adaptive learning in large populations
- Optimal training for adversarial games
- On the uniqueness of quantal response equilibria and its application to network games
- Tributes to Bill Sandholm
- Learning in nonatomic games. I: Finite action spaces and population games
- On Lyapunov functions and particle methods for regularized minimax problems
- On the robustness of learning in games with stochastically perturbed payoff observations
- Learning payoff functions in infinite games
- Learning strict Nash equilibria through reinforcement
- Exploration-exploitation in multi-agent learning: catastrophe theory meets game theory
- Aspiration-based reinforcement learning in repeated interaction games: An overview
- Reinforcement learning with restrictions on the action set
- Inertial game dynamics and applications to constrained optimization
- Solving for Best Responses and Equilibria in Extensive-Form Games with Reinforcement Learning Methods
- Reinforcement with fading memories
- scientific article; zbMATH DE number 5547959 (Why is no real title available?)
- The dynamics of generalized reinforcement learning
- Learning in network games
- On the convergence of gradient-like flows with noisy gradient input
- Cycles in adversarial regularized learning
- scientific article; zbMATH DE number 1931836 (Why is no real title available?)
- Continuous-time convergence rates in potential and monotone games
- Reinforcement Learning rules in a repeated game
- No-regret algorithms in on-line learning, games and convex optimization
- A unified stochastic approximation framework for learning in games
- Continuous time learning algorithms in optimization and game theory
- Q-learning in regularized mean-field games
- Memory loss can prevent chaos in games dynamics
- The rate of convergence of Bregman proximal methods: local geometry versus regularity versus sharpness
- Three-operator splitting for learning to predict equilibria in convex games
- Nested replicator dynamics, nested logit choice, and similarity-based learning
- Perturbed Bayesian best response dynamic in continuum games
- Regularized Bayesian best response learning in finite games
- Learning in random utility models via online decision problems
- Learning optimal strategies in a duel game
- Fast computation of optimal transport via entropy-regularized extragradient methods
- Mirror frameworks for relatively Lipschitz and monotone-like variational inequalities
- O(1/T) time-average convergence in a generalization of network zero-sum games via alternating gradient descent
- Risk-Sensitive Deep RL: Variance-Constrained Actor-Critic Provably Finds Globally Optimal Policy
- Replicator dynamics: old and new
- Learning in games using the imprecise Dirichlet model
This page was built for publication: Learning in games via reinforcement and regularization
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2833105)