A payoff-based policy gradient method in stochastic games with long-run average payoffs
From MaRDI portal
Cites work
- 10.1162/1532443041827880
- A one-measurement form of simultaneous perturbation stochastic approximation
- A Simple Adaptive Procedure Leading to Correlated Equilibrium
- A Stochastic Approximation Method
- A unified stochastic approximation framework for learning in games
- Brown's original fictitious play
- Existence and Uniqueness of Equilibrium Points for Concave N-Person Games
- Learning in games with continuous action sets and unknown payoff functions
- Learning Stationary Nash Equilibrium Policies in n-Player Stochastic Games with Independent Chains
- Mirror descent and nonlinear projected subgradient methods for convex optimization.
- Multi-agent reinforcement learning: a selective overview of theories and algorithms
- Online convex optimization in the bandit setting: gradient descent without a gradient
- Online learning and online convex optimization
- Online Markov decision processes
- Prediction, Learning, and Games
- Reinforcement learning. An introduction
- Solving variational inequalities with stochastic mirror-prox algorithm
- Stochastic Games
- Stochastic games
- V-learning -- a simple, efficient, decentralized algorithm for multiagent reinforcement learning
This page was built for publication: A payoff-based policy gradient method in stochastic games with long-run average payoffs
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6897257)