Reinforcement learning for individual optimal policy from heterogeneous data
From MaRDI portal
Cites work
- Batch policy learning in average reward Markov decision processes
- Distributed optimization and statistical learning via the alternating direction method of multipliers
- Estimating dynamic treatment regimes in mobile health using V-learning
- Estimating Optimal Infinite Horizon Dynamic Treatment Regimes via pT-Learning
- Estimation and Inference of Heterogeneous Treatment Effects using Random Forests
- Individualized Multidirectional Variable Selection
- Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
- Linear least-squares algorithms for temporal difference learning
- Nonparametric regression using deep neural networks with ReLU activation function
- Q-Learning with Linear Function Approximation
- Quasi-oracle estimation of heterogeneous treatment effects
- Reinforcement learning for individual optimal policy from heterogeneous data
- Weak convergence and empirical processes. With applications to statistics
Cited in
(2)
This page was built for publication: Reinforcement learning for individual optimal policy from heterogeneous data
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6926373)