Reward Biased Maximum Likelihood Estimation for Learning in Constrained MDPs
From MaRDI portal
Abstract: We use the Reward Biased Maximum Likelihood Estimation (RBMLE) algorithm to learn optimal policies for constrained Markov Decision Processes (CMDPs). We analyze the learning regrets of RBMLE.
This page was built for publication: Reward Biased Maximum Likelihood Estimation for Learning in Constrained MDPs
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6368817)