Partially observable RL: benign structures and simple generic algorithms
From MaRDI portal
Cites work
- scientific article; zbMATH DE number 1420699 (Why is no real title available?)
- scientific article; zbMATH DE number 3233996 (Why is no real title available?)
- 10.1162/153244303321897663
- 10.1162/153244303765208377
- A Decomposition for Three-Way Arrays
- A spectral algorithm for learning hidden Markov models
- Bandit algorithms
- Complexity of finite-horizon Markov decision process problems
- Deep Blue
- Learning nonsingular phylogenies and hidden Markov models
- Near-optimal regret bounds for reinforcement learning
- On the Computational Complexity of Stochastic Controller Optimization in POMDPs
- Optimal control of Markov processes with incomplete state information
- Optimistic MLE: a generic model-based algorithm for partially observable sequential decision making
- Revisiting Ho–Kalman-Based System Identification: Robustness and Finite-Sample Analysis
- Superhuman AI for multiplayer poker
- Tensor decompositions for learning latent variable models
- The Complexity of Decentralized Control of Markov Decision Processes
- The Complexity of Markov Decision Processes
- Unified algorithms for RL with decision-estimation coefficients: PAC, reward-free, preference-based learning and beyond
- V-learning -- a simple, efficient, decentralized algorithm for multiagent reinforcement learning
- Zur Theorie der Gesellschaftsspiele.
This page was built for publication: Partially observable RL: benign structures and simple generic algorithms
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6860954)