Multi-agent reinforcement learning: a selective overview of theories and algorithms
From MaRDI portal
Publication:2094040
Cites work
- ${{\cal Q} {\cal D}}$-Learning: A Collaborative Distributed Strategy for Multi-Agent Reinforcement Learning Through ${\rm Consensus} + {\rm Innovations}$
- 10.1162/153244303765208377
- 10.1162/1532443041827880
- \(H^ \infty\)-optimal control and related minimax design problems. A dynamic game approach.
- \({\mathcal Q}\)-learning
- A comprehensive survey on safe reinforcement learning
- A concise introduction to decentralized POMDPs
- A course in game theory.
- A Distributed Actor-Critic Algorithm and Applications to Mobile Sensor Network Coordination Problems
- A general class of adaptive strategies
- A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
- A near-optimal polynomial time algorithm for learning in certain classes of stochastic games
- A Simple Adaptive Procedure Leading to Correlated Equilibrium
- Achieving Geometric Convergence for Distributed Optimization Over Time-Varying Graphs
- Adaptive game playing using multiplicative weights
- Algorithms for discounted stochastic games
- An Adaptive Sampling Algorithm for Solving Markov Decision Processes
- An emphatic approach to the problem of off-policy temporal-difference learning
- An iterative method of solving a game
- Analysis of Hannan consistent selection for Monte Carlo tree search in simultaneous move games
- Approximate Markov-Nash equilibria for discrete-time risk-sensitive mean-field games
- Approximate Nash equilibria in partially observed stochastic games with mean-field interactions
- AWESOME: a general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents
- Belief affirming in learning processes
- Consistency and cautious fictitious play
- Consistency of vanishingly smooth fictitious play
- Convergence results for single-step on-policy reinforcement-learning algorithms
- Cooperative Convex Optimization in Networked Systems: Augmented Lagrangian Algorithms With Directed Gossip Communication
- Decentralized Q-Learning for Stochastic Teams and Games
- Decentralized Stochastic Control with Partial History Sharing: A Common Information Approach
- Decomposition of dynamic team decision problems
- DeepStack: expert-level artificial intelligence in heads-up no-limit poker
- Diffusion Strategies Outperform Consensus Strategies for Distributed Estimation Over Adaptive Networks
- Discounted Markov games: Generalized policy iteration method
- Discrete-time average-cost mean-field games on Polish spaces
- Discrete-time stochastic control and dynamic potential games. The Euler-equation approach
- Distributed learning and cooperative control for multi-agent systems
- Distributed learning of average belief over networks using sequential observations
- Distributed Policy Evaluation Under Multiple Behavior Strategies
- Distributed Stochastic Approximation: Weak Convergence and Network Design
- Distributed Subgradient Methods for Multi-Agent Optimization
- Dynamic Potential Games With Constraints: Fundamentals and Applications in Communications
- Dynamic programming and optimal control. Vol. 1.
- Efficient computation of behavior strategies
- Efficient computation of equilibria for extensive two-person games
- Expected policy gradients for reinforcement learning
- Fast algorithms for finding randomized strategies in game trees
- Finite mean field games: fictitious play and convergence to a first order continuous mean field game
- Finite-time analysis of the multiarmed bandit problem
- Finite-time bounds for fitted value iteration
- Finite-time performance of distributed temporal-difference learning with linear function approximation
- Generalised weakened fictitious play
- Global convergence of policy gradient methods to (almost) locally optimal policies
- Handbook of dynamic game theory. In 2 volumes
- Harnessing Smoothness to Accelerate Distributed Optimization
- scientific article; zbMATH DE number 5957269 (Why is no real title available?)
- scientific article; zbMATH DE number 3128728 (Why is no real title available?)
- scientific article; zbMATH DE number 5145289 (Why is no real title available?)
- scientific article; zbMATH DE number 5547972 (Why is no real title available?)
- scientific article; zbMATH DE number 125484 (Why is no real title available?)
- scientific article; zbMATH DE number 1243371 (Why is no real title available?)
- scientific article; zbMATH DE number 1134975 (Why is no real title available?)
- scientific article; zbMATH DE number 1972910 (Why is no real title available?)
- scientific article; zbMATH DE number 1509479 (Why is no real title available?)
- scientific article; zbMATH DE number 1753152 (Why is no real title available?)
- scientific article; zbMATH DE number 1795161 (Why is no real title available?)
- scientific article; zbMATH DE number 7064064 (Why is no real title available?)
- scientific article; zbMATH DE number 3062455 (Why is no real title available?)
- scientific article; zbMATH DE number 3069635 (Why is no real title available?)
- scientific article; zbMATH DE number 3078991 (Why is no real title available?)
- If multi-agent learning is the answer, what is the question?
- Large population stochastic dynamic games: closed-loop McKean-Vlasov systems and the Nash certainty equivalence principle
- Learning Near-Optimal Policies with Bellman-Residual Minimization Based Fitted Policy Iteration and a Single Sample Path
- Learning, regret minimization, and equilibria
- Markov-Nash equilibria in mean-field games with discounted cost
- Mean field games
- Mean field games and mean field type control theory
- Minimizing finite sums with the stochastic average gradient
- Model-free Q-learning designs for linear discrete-time zero-sum games with application to H^ control
- Monotonic value function factorisation for deep multi-agent reinforcement learning
- Multiagent learning using a variable learning rate
- Multiagent Systems
- Natural actor-critic algorithms
- Near-optimal regret bounds for reinforcement learning
- No-regret dynamics and fictitious play
- On Nonterminating Stochastic Games
- On the Global Convergence of Stochastic Fictitious Play
- OnActor-Critic Algorithms
- Optimal Decentralized Control of Coupled Subsystems With Control Sharing
- Optimal stochastic linear systems with exponential performance criteria and their relation to deterministic differential games
- Optimally solving Dec-POMDPs as continuous-state MDPs
- Performance Bounds in $L_p$‐norm for Approximate Value Iteration
- Policy evaluation with temporal differences: a survey and comparison
- Potential games
- Prediction, Learning, and Games
- Regularized policy iteration with nonparametric function spaces
- Reinforcement learning with replacing eligibility traces
- Reinforcement learning. An introduction
- Revisiting CFR^+ and alternating updates
- Risk-Sensitive Mean-Field Games
- Sample mean based index policies by O(log n) regret for the multi-armed bandit problem
- Sampled fictitious play is Hannan consistent
- Settling the complexity of computing two-player Nash equilibria
- Simple statistical gradient-following algorithms for connectionist reinforcement learning
- State of the Art—A Survey of Partially Observable Markov Decision Processes: Theory, Models, and Algorithms
- Stochastic approximation. A dynamical systems viewpoint.
- Stochastic Approximations and Differential Inclusions
- Stochastic Games
- Stochastic networked control systems. Stabilization and optimization under information constraints
- Stochastic Proximal Gradient Consensus Over Random Networks
- Strategy iteration is strongly polynomial for 2-player turn-based stochastic games with a constant discount factor
- Subjectivity and correlation in randomized strategies
- Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
- Superhuman AI for multiplayer poker
- The challenge of poker
- The complexity of computing a Nash equilibrium
- The Complexity of Decentralized Control of Markov Decision Processes
- The complexity of two-person zero-sum games in extensive form
- The Evolution of Conventions
- The Nonstochastic Multiarmed Bandit Problem
- The weighted majority algorithm
- Value iteration algorithm for mean-field games
Cited in
(59)- The possible and the impossible in multi-agent learning
- Multi-agent reinforcement learning algorithm to solve a partially-observable multi-agent problem in disaster response
- Fully asynchronous policy evaluation in distributed reinforcement learning over networks
- Stackelberg population dynamics: a predictive-sensitivity approach
- A mini review on UAV mission planning
- Dynamics and risk sharing in groups of selfish individuals
- scientific article; zbMATH DE number 2086990 (Why is no real title available?)
- Mean-field controls with Q-learning for cooperative MARL: convergence and complexity analysis
- Scalable Reinforcement Learning for Multiagent Networked Systems
- Scalable Online Planning for Multi-Agent MDPs
- Fictitious play in zero-sum stochastic games
- Toward multi-target self-organizing pursuit in a partially observable Markov game
- Zeroth-order algorithms for nonconvex-strongly-concave minimax problems with improved complexities
- TEAMSTER: model-based reinforcement learning for ad hoc teamwork
- A multiagent reinforcement learning framework for off-policy evaluation in two-sided markets
- Entropy regularized actor-critic based multi-agent deep reinforcement learning for stochastic games
- Learning Stationary Nash Equilibrium Policies in n-Player Stochastic Games with Independent Chains
- Multi-agent natural actor-critic reinforcement learning algorithms
- Robustness and sample complexity of model-based MARL for general-sum Markov games
- Approximated multi-agent fitted Q iteration
- Independent learning in stochastic games
- Reinforcement learning in a prisoner's dilemma
- Finite-time error bounds for distributed linear stochastic approximation
- An optimal Bayesian intervention policy in response to unknown dynamic cell stimuli
- Predator-prey survival pressure is sufficient to evolve swarming behaviors
- Cournot policy model: rethinking centralized training in multi-agent reinforcement learning
- On neural networks application in integral sliding mode control
- Statistical inference for generative adversarial networks and other minimax problems
- Recent developments in machine learning methods for stochastic control and games
- Convergence properties of gradient-based methods for minimax problems with nonlinear constraints
- Team variance optimization of n-player stochastic games with separately controlled chains
- Partially observable multiagent reinforcement learning with information sharing
- Multi-agent game strategies for kill chain optimization in networked aerial combat systems
- Distributed policy gradient with variance reduction in multi-agent reinforcement learning
- Preference-based opponent shaping in differentiable games
- A nonzero-sum game with reinforcement learning under mean-variance framework
- Localized multi-agent reinforcement learning for cooperative management of supply chains
- A passivity analysis for nonlinear consensus on balanced digraphs
- Learning with linear function approximations in mean-field control
- A payoff-based policy gradient method in stochastic games with long-run average payoffs
- Regularized minimax-V learning for solving randomly terminating two-player zero-sum Markov games
- MF-OML: online mean-field reinforcement learning with occupation measures for large population games
- Provably efficient information-directed sampling algorithms for multi-agent reinforcement learning
- Logit-Q dynamics for efficient learning in stochastic teams
- Cluster approximate synchronization for probabilistic asynchronous finite field networks
- A computational stochastic dynamic model to assess the risk of breakup in a romantic relationship
- Mean-field games with finitely many players: independent learning and subjectivity
- Balancing individual and collective strategies: a new approach in metaheuristic optimization
- Grounded predictions of teamwork as a one-shot game: a multiagent multi-armed bandits approach
- Entropic mean-field min-max problems via best response flow
- Learning collusive strategies with function approximation algorithms
- Enhancing multi-agent deep reinforcement learning for flexible job-shop scheduling through constraint programming
- Dynamic repair and maintenance of heterogeneous machines dispersed on a network: a rollout method for online reinforcement learning
- Synthesising reward machines for cooperative multi-agent reinforcement learning
- Reinforcement learning for infinite-dimensional systems
- Client selection for federated policy optimization with environment heterogeneity
- Homing through reinforcement learning
- Multi-agent reinforcement learning control with Lyapunov stability guarantees for cooperative systems
- Evolutionary Stability Does Not Guarantee Learning Accessibility: A Multi-Agent Reinforcement Learning Perspective on Cooperation Emergence
This page was built for publication: Multi-agent reinforcement learning: a selective overview of theories and algorithms
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2094040)