Policy space identification in configurable environments
From MaRDI portal
Recommendations
- Identification of optimal policies in Markov decision processes
- Stochastic policy design in a learning environment with rational expectations.
- Interactive policy learning through confidence-based autonomy
- k-Certainty Exploration Method: an action selector to identify the environment in reinforcement learning
- Policy iterations for reinforcement learning problems in continuous time and space -- fundamental theory and methods
Cites work
- A tail inequality for quadratic forms of subgaussian random vectors
- Concentration inequalities. A nonasymptotic theory of independence
- Finite-sample analysis of least-squares policy iteration
- Generalized inverses. Theory and applications.
- scientific article; zbMATH DE number 3146392 (Why is no real title available?)
- scientific article; zbMATH DE number 3173999 (Why is no real title available?)
- scientific article; zbMATH DE number 44577 (Why is no real title available?)
- scientific article; zbMATH DE number 46578 (Why is no real title available?)
- scientific article; zbMATH DE number 700091 (Why is no real title available?)
- Identification in Parametric Models
- Importance sampling techniques for policy optimization
- Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
- Nonasymptotic sequential tests for overlapping hypotheses applied to near-optimal arm identification in bandit models
- Rates of convergence for empirical processes of stationary mixing sequences
- Reinforcement learning. An introduction
- The Large-Sample Distribution of the Likelihood Ratio for Testing Composite Hypotheses
- The philosophy of Bayes factors and the quantification of statistical evidence
This page was built for publication: Policy space identification in configurable environments
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2163245)