Reinforcement Learning with Continuous Actions Under Unmeasured Confounding
From MaRDI portal
Cites work
- A semiparametric instrumental variable approach to optimal treatment regimes under endogeneity
- Batch policy learning in average reward Markov decision processes
- Breaking the curse of dimensionality with convex neural networks
- Estimating dynamic treatment regimes in mobile health using V-learning
- Estimating Optimal Infinite Horizon Dynamic Treatment Regimes via pT-Learning
- scientific article; zbMATH DE number 1420699 (Why is no real title available?)
- Identifying causal effects with proxy variables of an unmeasured confounder
- Instrumental Variable Estimation of Nonparametric Models
- Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
- Linear integral equations
- Off-Policy Confidence Interval Estimation with Confounded Markov Decision Process
- Off-policy estimation of long-term average outcomes with applications to mobile health
- On the completeness condition in nonparametric instrumental problems
- Proximal Learning for Individualized Treatment Regimes Under Unmeasured Confounding
- Randomized sketches for kernels: fast and optimal nonparametric regression
- Regularized policy iteration with nonparametric function spaces
- Reinforcement learning. An introduction
- Semiparametric Proximal Causal Inference
- Sensitivity Analysis via the Proportion of Unmeasured Confounding
- Statistical inference of the value function for reinforcement learning in infinite-horizon settings
This page was built for publication: Reinforcement Learning with Continuous Actions Under Unmeasured Confounding
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q7231146)