Markov decision processes with arbitrary reward processes
From MaRDI portal
Recommendations
- Online Markov decision processes
- An Online Policy Gradient Algorithm for Markov Decision Processes with Continuous States and Actions
- Learning algorithms for Markov decision processes
- Adaptive aggregation for reinforcement learning in average reward Markov decision processes
- Online regret bounds for Markov decision processes with deterministic transitions
Cited in
(30)- Approachability in Stackelberg stochastic games with vector costs
- Reinforcement learning with immediate rewards and linear hypotheses
- Batch policy learning in average reward Markov decision processes
- Efficient PAC learning for episodic tasks with acyclic state spaces
- Reinforcement learning in robust Markov decision processes
- Online learning in Markov decision processes with continuous actions
- A learning algorithm for communicating Markov decision processes with unknown transition matrices
- 10.1162/153244303765208377
- 10.1162/153244303768966148
- Making virtual learning environment more intelligent: an application of Markov decision process
- Value function based reinforcement learning in changing Markovian environments
- Trading value and information in mdps
- Online Markov decision processes
- Long-Run Rewards for Markov Automata
- Asymptotic Learnability of Reinforcement Problems with Arbitrary Dependence
- Online Regret Bounds for Markov Decision Processes with Deterministic Transitions
- scientific article; zbMATH DE number 5547823 (Why is no real title available?)
- Adaptive aggregation for reinforcement learning in average reward Markov decision processes
- Bayesian learning of noisy Markov decision processes
- scientific article; zbMATH DE number 1931851 (Why is no real title available?)
- Reversible Markov Decision Processes with an Average-Reward Criterion
- Game of thrones: fully distributed learning for multiplayer bandits
- Online learning over a finite action set with limited switching
- Learning and planning for time-varying MDPs using maximum likelihood estimation
- An Online Policy Gradient Algorithm for Markov Decision Processes with Continuous States and Actions
- A holistic matrix norm-based alternative solution method for Markov reward games
- Model-free reinforcement learning for branching Markov decision processes
- Markovian decision programming with recursive vector-reward
- On the possibility of learning in reactive environments with arbitrary dependence
- Online regret bounds for Markov decision processes with deterministic transitions
This page was built for publication: Markov decision processes with arbitrary reward processes
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3169064)