Two classes Markov decision processes with perturbations
Two classes of Markov decision processes (MDP) (with denumerable state space \(S\) and denumerable action set \(A)\) with perturbations are discussed. Under the definition of a \(\delta\)-optimal policy, a perturbation model \(P_\varepsilon(D)\) for the discrete-time non-stationary MDP with respect to a maximization criterion of limiting average expected reward, and a perturbation model \(C_\varepsilon(D)\) for the continuous-time stationary MDP with respect to the criteria of maximization of the discounted expected reward are proposed where the transition probabilities \((p_{n+1} (j|i,a) (\varepsilon)\) and \(q(j|i,a) (\varepsilon)\), \(i,j\in S\), \(a\in A)\) in the two cases are taken as perturbed according to a so-called disturbance set \(D\) [cf. \textit{M. Abbad} and \textit{J. A. Filar}, IEEE Trans. Autom. Control 37, 1415-1420 (1992; Zbl 0763.90091)]. It is then proved that if \(\pi\) is an optimal stochastic (Markov or stationary) policy before perturbation, then for any \(\delta>0\) there exists an \(\varepsilon >0\) such that \(\pi\) is \(\delta\)-optimal in the perturbation model \((P_\varepsilon (D)\) or \(C_\varepsilon(D)\), respectively).
- Singulary perturbed Markov control problem: Limiting average cost
- Weighted discounted Markov decision processes with perturbation
- Semi-infinite weighted Markov decision processes with perturbation.
- Estimates for perturbations of general discounted Markov control chains
- Weighted Markov decision processes with perturbation
- Nonuniqueness versus uniqueness of optimal policies in convex discounted Markov decision processes
- Perturbation and stability theory for Markov control problems
- scientific article; zbMATH DE number 1944273
- scientific article; zbMATH DE number 2159039
- Weighted Markov decision processes with perturbation
- From perturbation analysis to Markov decision processes and reinforcement learning
- scientific article; zbMATH DE number 3987072 (Why is no real title available?)
- Perturbation and stability theory for Markov control problems
- scientific article; zbMATH DE number 2098845 (Why is no real title available?)
- Estimates for perturbations of average Markov decision processes with a minimal state and upper bounded by stochastically ordered Markov chains.
This page was built for publication: Two classes Markov decision processes with perturbations
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2739190)