An emphatic approach to the problem of off-policy temporal-difference learning
From MaRDI portal
(Redirected from Publication:2810885)
Recommendations
Cited in
(13)- Adaptive importance sampling for value function approximation in off-policy reinforcement learning
- Off-policy temporal difference learning with distribution adaptation in fast mixing chains
- Multi-agent reinforcement learning: a selective overview of theories and algorithms
- \(\text{Q}(\lambda)\) with off-policy corrections
- Policy evaluation with temporal differences: a survey and comparison
- Statistical inference for online decision making via stochastic gradient descent
- Efficiently Breaking the Curse of Horizon in Off-Policy Evaluation with Double Reinforcement Learning
- Gradient temporal-difference learning for off-policy evaluation using emphatic weightings
- Estimating Optimal Infinite Horizon Dynamic Treatment Regimes via pT-Learning
- Distributed consensus-based multi-agent temporal-difference learning
- Online Bootstrap Inference For Policy Evaluation In Reinforcement Learning
- An approximate policy iteration viewpoint of actor-critic algorithms
- The ODE method for stochastic approximation and reinforcement learning with Markovian noise
This page was built for publication: An emphatic approach to the problem of off-policy temporal-difference learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2810885)