Efficiently Breaking the Curse of Horizon in Off-Policy Evaluation with Double Reinforcement Learning
From MaRDI portal
(Redirected from Publication:5060503)
Recommendations
- Double reinforcement learning for efficient off-policy evaluation in Markov decision processes
- Doubly robust policy evaluation and optimization
- scientific article; zbMATH DE number 1753153
- Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning
- Policy learning for time-bounded reachability in continuous-time Markov decision processes via doubly-stochastic gradient ascent
- An emphatic approach to the problem of off-policy temporal-difference learning
- scientific article; zbMATH DE number 1753152
- Off-policy linear temporal difference learning algorithms with a generalized oblique projection
- Reinforcement learning in sparse-reward environments with hindsight policy gradients
Cites work
- 10.1162/1532443041827907
- Adjusting for Nonignorable Drop-Out Using Semiparametric Nonresponse Models
- Asymptotic Statistics
- Basic properties of strong mixing conditions. A survey and some open questions
- Characterization of parameters with a mixed bias property
- Comment: Understanding OR, PS and DR
- Consistent estimation of the influence function of locally asymptotically linear estimators
- Double reinforcement learning for efficient off-policy evaluation in Markov decision processes
- Double/debiased machine learning for treatment and structural parameters
- Doubly robust policy evaluation and optimization
- Dynamic programming and optimal control. Vol. 2
- Efficient Estimation of Average Treatment Effects Using the Estimated Propensity Score
- Efficient estimation of panel data models with sequential moment restrictions
- Estimating dynamic treatment regimes in mobile health using V-learning
- Estimation of Regression Coefficients When Some Regressors Are Not Always Observed
- Generalized TD learning
- scientific article; zbMATH DE number 1181283 (Why is no real title available?)
- scientific article; zbMATH DE number 3349105 (Why is no real title available?)
- Introduction to empirical processes and semiparametric inference
- Irregular identification, support conditions, and inverse weight estimation
- Least squares policy evaluation algorithms with linear function approximation
- Least squares temporal difference methods: An analysis under general conditions
- Marginal Mean Models for Dynamic Regimes
- Markov Chains and Stochastic Stability
- On the Markov chain central limit theorem
- Optimal Dynamic Treatment Regimes
- Reinforcement learning. An introduction
- Semiparametric efficiency bounds
- Semiparametric theory and missing data.
- Sieve Extremum Estimates for Weakly Dependent Data
Cited in
(20)- Double reinforcement learning for efficient off-policy evaluation in Markov decision processes
- A multiagent reinforcement learning framework for off-policy evaluation in two-sided markets
- Off-policy evaluation in partially observed Markov decision processes under sequential ignorability
- Projected state-action balancing weights for offline reinforcement learning
- Online Bootstrap Inference For Policy Evaluation In Reinforcement Learning
- Off-policy evaluation for tabular reinforcement learning with synthetic trajectories
- Deep spectral Q-learning with application to mobile health
- Predicting and optimizing marketing performance in dynamic markets
- Reliable off-policy evaluation for reinforcement learning
- Proximal reinforcement learning: efficient off-policy evaluation in partially observed Markov decision processes
- On the statistical complexity for offline and low-adaptive reinforcement learning with structures
- Asymptotic inference for multi-stage stationary treatment policy with variable selection
- Distributional Off-Policy Evaluation in Reinforcement Learning
- Graphical criteria for the identification of marginal causal effects in continuous-time survival and event-history analyses
- Randomization inference when N equals one
- Off-Policy Evaluation in Doubly Inhomogeneous Environments
- Multivariate dynamic mediation analysis under a reinforcement learning framework
- A review of causal decision making
- A review of off-policy evaluation in reinforcement learning
- Risk-Sensitive Deep RL: Variance-Constrained Actor-Critic Provably Finds Globally Optimal Policy
This page was built for publication: Efficiently Breaking the Curse of Horizon in Off-Policy Evaluation with Double Reinforcement Learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5060503)