Simulation-based search
Summary: Planning is one of the oldest and most important problems in artificial intelligence. Simulation-based search algorithms, such as AlphaZero, have achieved superhuman performance in chess and Go and are used widely in real-world applications of planning. In this paper we provide a unified framework for simulation-based search. Algorithms in this framework interleave operators for policy evaluation (better estimating the value function of the current policy) and policy improvement (using the value function to form a better policy). These operators are applied to states and actions that are sampled in sequential trajectories, and that may branch recursively into other sampled trajectories. The value function and policy may also be represented by a function approximator. Our framework includes a broad family of search algorithms that includes Monte-Carlo tree search, sparse sampling, nested Monte-Carlo search, classification-based policy iteration, and AlphaZero. For the entire collection see [Zbl 07816360].
- On Monte Carlo tree search and reinforcement learning
- Temporal-difference search in Computer Go
- Analyzing simulations in Monte-Carlo tree search for the game of Go
- Sample-based tree search with fixed and adaptive state abstractions
- A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
- A comparison of minimax tree search algorithms
- A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
- A sparse sampling algorithm for near-optimal planning in large Markov decision processes
- Approximate dynamic programming. Solving the curses of dimensionality
- Convergence results for single-step on-policy reinforcement-learning algorithms
- Deep Blue
- Dynamic programming and optimal control. Vol. 1.
- Finite-time analysis of the multiarmed bandit problem
- scientific article; zbMATH DE number 700091 (Why is no real title available?)
- scientific article; zbMATH DE number 6542806 (Why is no real title available?)
- Model predictive control: Theory and practice - a survey
- Temporal-difference search in Computer Go
- World-championship-caliber Scrabble*
This page was built for publication: Simulation-based search
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6198646)