An analysis of temporal-difference learning with function approximation
From MaRDI portal
Recommendations
- On the convergence of temporal-difference learning with linear function approximation
- Average cost temporal-difference learning
- Least squares policy evaluation algorithms with linear function approximation
- Asymptotic analysis of temporal-difference learning algorithms with constant step-sizes
- Asymptotic analysis of temporal-difference learning algorithms with constant step-sizes
Cited in
(only showing first 100 items - show all)- A formal framework and extensions for function approximation in learning classifier systems
- Projected equation methods for approximate solution of large linear systems
- Reinforcement distribution in fuzzy Q-learning
- Natural actor-critic algorithms
- Rationality and intelligence
- A reinforcement learning adaptive fuzzy controller for robots.
- On the existence of fixed points for approximate value iteration and temporal-difference learning
- An incremental off-policy search in a model-free Markov decision process using a single sample path
- An online prediction algorithm for reinforcement learning with linear function approximation using cross entropy method
- Variance-constrained actor-critic algorithms for discounted and average reward MDPs
- Real-time reinforcement learning by sequential actor-critics and experience replay
- Off-policy temporal difference learning with distribution adaptation in fast mixing chains
- Energy contracts management by stochastic programming techniques
- Concentration bounds for temporal difference learning with linear function approximation: the case of batch data and uniform sampling
- Transmission scheduling for multi-process multi-sensor remote estimation via approximate dynamic programming
- Deep reinforcement learning for inventory control: a roadmap
- Fundamental design principles for reinforcement learning algorithms
- Neural circuits for learning context-dependent associations of stimuli
- A Q-learning predictive control scheme with guaranteed stability
- A review on deep reinforcement learning for fluid mechanics
- Restricted gradient-descent algorithm for value-function approximation in reinforcement learning
- Basis function adaptation in temporal difference reinforcement learning
- An actor-critic algorithm for constrained Markov decision processes
- The single-node dynamic service scheduling and dispatching problem
- Reinforcement learning based algorithms for average cost Markov decision processes
- An approximate dynamic programming approach to the admission control of elective patients
- Adaptive critic design with graph Laplacian for online learning control of nonlinear systems
- Chaotic dynamics and convergence analysis of temporal difference algorithms with bang-bang control
- Flow shop scheduling with reinforcement learning
- Approximate policy iteration: a survey and some new methods
- A review of stochastic algorithms with continuous value function approximation and some new approximate policy iteration algorithms for multidimensional continuous applications
- Adaptive importance sampling for control and inference
- Solving average cost Markov decision processes by means of a two-phase time aggregation algorithm
- Bias and variance approximation in value function estimates
- Multiscale Q-learning with linear function approximation
- Perspectives of approximate dynamic programming
- Variance regularization in sequential Bayesian optimization
- Hybrid MDP based integrated hierarchical Q-learning
- On the Asymptotic Equivalence Between Differential Hebbian and Temporal Difference Learning
- Robust reinforcement learning control with static and dynamic stability
- Parallel dynamic water supply scheduling in a cluster of computers
- Asymptotic analysis of value prediction by well-specified and misspecified models
- From infinite to finite programs: explicit error bounds with applications to approximate dynamic programming
- On-policy concurrent reinforcement learning
- Temporal difference-based policy iteration for optimal control of stochastic systems
- A Sarsa() algorithm based on double-layer fuzzy reasoning
- Least squares temporal difference methods: An analysis under general conditions
- Bayesian exploration for approximate dynamic programming
- Risk-averse learning by temporal difference methods with Markov risk measures
- Finite-time performance of distributed temporal-difference learning with linear function approximation
- A finite time analysis of temporal difference learning with linear function approximation
- High-order fully actuated system approaches. VIII: Optimal control with application in spacecraft attitude stabilisation
- Simple and optimal methods for stochastic variational inequalities. II: Markovian noise and policy evaluation in reinforcement learning
- Stochastic recursive inclusions with non-additive iterate-dependent Markov noise
- Risk-Sensitive Reinforcement Learning via Policy Gradient Search
- Full gradient DQN reinforcement learning: a provably convergent scheme
- Is Temporal Difference Learning Optimal? An Instance-Dependent Analysis
- A tutorial on linear function approximators for dynamic programming and reinforcement learning
- Continuous-time robust dynamic programming
- Deep exploration via randomized value functions
- Actor-critic algorithms with online feature adaptation
- Toward nonlinear local reinforcement learning rules through neuroevolution
- Quadratic approximate dynamic programming for input-affine systems
- Q-Learning with Linear Function Approximation
- Relational Sequence Learning
- The Borkar-Meyn theorem for asynchronous stochastic approximations
- Concentration of Contractive Stochastic Approximation and Reinforcement Learning
- Least squares policy iteration with instrumental variables vs. direct policy search: comparison against optimal benchmarks using energy storage
- Accelerated and Instance-Optimal Policy Evaluation with Linear Function Approximation
- Asymptotic analysis of temporal-difference learning algorithms with constant step-sizes
- Stochastic approximation
- Asymptotic analysis of temporal-difference learning algorithms with constant step-sizes
- On the convergence of temporal-difference learning with linear function approximation
- Stochastic approximation algorithms: overview and recent trends.
- A Small Gain Analysis of Single Timescale Actor Critic
- Finite-time convergence rates of distributed local stochastic approximation
- Convergence of stochastic approximation via martingale and converse Lyapunov methods
- A Lyapunov-based version of the value iteration algorithm formulated as a discrete-time switched affine system
- On the sample complexity of actor-critic method for reinforcement learning with function approximation
- Approximate Q Learning for Controlled Diffusion Processes and Its Near Optimality
- Uncovering instabilities in variational-quantum deep Q-networks
- Gradient temporal-difference learning for off-policy evaluation using emphatic weightings
- Target Network and Truncation Overcome the Deadly Triad in \(\boldsymbol{Q}\)-Learning
- From Reinforcement Learning to Deep Reinforcement Learning: An Overview
- An actor-critic algorithm with function approximation for discounted cost constrained Markov decision processes
- Premium control with reinforcement learning
- Online Bootstrap Inference For Policy Evaluation In Reinforcement Learning
- Eligibility traces and forgetting factor in recursive least-squares-based temporal difference
- Finite-time error bounds for distributed linear stochastic approximation
- Convergence of entropy-regularized natural policy gradient with linear function approximation
- A functional model method for nonconvex nonsmooth conditional stochastic optimization
- Convergence of stochastic approximation via martingale and converse Lyapunov methods
- Optimal policy evaluation using kernel-based temporal difference methods
- Stochastic approximation and reinforcement learning: the interface and a little beyond
- The central role of the loss function in reinforcement learning
- Online estimation and inference for robust policy evaluation in reinforcement learning
- An approximate policy iteration viewpoint of actor-critic algorithms
- Optimal consensus control of nonlinear multi-agent systems by data-driven optimistic policy iteration
- High-precision geosteering via reinforcement learning and particle filters
- Proximal algorithms and temporal difference methods for solving fixed point problems
This page was built for publication: An analysis of temporal-difference learning with function approximation
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4362297)