Fundamental design principles for reinforcement learning algorithms
From MaRDI portal
Publication:2094028
Cites work
- \({\mathcal Q}\)-learning
- A concentration bound for stochastic approximation via Alekseev's formula
- A finite time analysis of temporal difference learning with linear function approximation
- A generalized Kalman filter for fixed point approximation and efficient temporal-difference learning
- A Newton-Raphson version of the multivariate Robbins-Monro procedure
- A Stochastic Approximation Method
- A Tutorial on Thompson Sampling
- Acceleration of Stochastic Approximation by Averaging
- Algorithms for reinforcement learning.
- An analysis of temporal-difference learning with function approximation
- An Extension of the Robbins-Monro Procedure
- Applications of a Kushner and Clark lemma to general classes of stochastic algorithms
- Asynchronous stochastic approximation and Q-learning
- Average cost temporal-difference learning
- Bandit algorithms
- Computable bounds for geometric convergence rates of Markov chains
- Computable exponential convergence rates for stochastically ordered Markov processes
- Concentration inequalities for Markov chains by Marton couplings and spectral methods
- Control Techniques for Complex Networks
- Convergence rate of linear two-time-scale stochastic approximation.
- Deep exploration via randomized value functions
- Dynamic programming and optimal control. Vol. 2
- Hoeffding's inequality for uniformly ergodic Markov chains
- scientific article; zbMATH DE number 5957196 (Why is no real title available?)
- scientific article; zbMATH DE number 48727 (Why is no real title available?)
- scientific article; zbMATH DE number 1043533 (Why is no real title available?)
- scientific article; zbMATH DE number 1158743 (Why is no real title available?)
- Large deviation asymptotics and control variates for simulating large functions
- Large deviation asymptotics for busy periods
- Learning algorithms for Markov decision processes with average cost
- Markov Chains and Stochastic Stability
- Multidimensional Stochastic Approximation Methods
- New method of stochastic approximation type
- Oja's algorithm for graph clustering, Markov spectral decomposition, and risk sensitive control
- On a Stochastic Approximation Method
- OnActor-Critic Algorithms
- Optimal stopping of Markov processes: Hilbert space theory, approximation algorithms, and an application to pricing high-dimensional financial derivatives
- Q-learning and policy iteration algorithms for stochastic shortest path problems
- Reinforcement learning. An introduction
- Spectral theory and limit theorems for geometrically ergodic Markov processes
- Stochastic approximation. A dynamical systems viewpoint.
- Stochastic approximations for finite-state Markov chains
- Stochastic Estimation of the Maximum of a Regression Function
- The O.D.E. Method for Convergence of Stochastic Approximation and Reinforcement Learning
This page was built for publication: Fundamental design principles for reinforcement learning algorithms
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2094028)