Learning algorithms for Markov decision processes with average cost
From MaRDI portal
Recommendations
- Reinforcement learning based algorithms for average cost Markov decision processes
- Empirical Q-value iteration
- Reinforcement learning for long-run average cost.
- Actor-Critic--Type Learning Algorithms for Markov Decision Processes
- Q-learning and policy iteration algorithms for stochastic shortest path problems
Cited in
(42)- \(L^\ast\)-based learning of Markov decision processes (extended version)
- Opportunistic Transmission over Randomly Varying Channels
- Batch policy learning in average reward Markov decision processes
- Risk-Sensitive Reinforcement Learning via Policy Gradient Search
- Q-learning and enhanced policy iteration in discounted dynamic programming
- Look-ahead control of conveyor-serviced production station by using potential-based online policy iteration
- Whittle index based Q-learning for restless bandits with average reward
- The ODE method for stochastic approximation and reinforcement learning with Markovian noise
- Optimal sensor scheduling for remote state estimation with limited bandwidth: a deep reinforcement learning approach
- Learning algorithms for Markov decision processes
- Reinforcement learning for long-run average cost.
- Relative value iteration algorithm with soft state aggregation
- Solutions of the average cost optimality equation for Markov decision processes with weakly continuous kernel: the fixed-point approach revisited
- Approximation of average cost Markov decision processes using empirical distributions and concentration inequalities
- Stochastic Fixed-Point Iterations for Nonexpansive Maps: Convergence and Error Bounds
- Natural actor-critic algorithms
- Multiscale Q-learning with linear function approximation
- A reinforcement learning algorithm based on policy iteration for average reward: Empirical results with yield management and convergence analysis
- A sojourn-based approach to semi-Markov reinforcement learning
- Scalable Whittle index policy for real-time storage allocation in railway container yard
- Variance-penalized Markov decision processes: dynamic programming and reinforcement learning techniques
- A perturbation approach to approximate value iteration for average cost Markov decision processes with Borel spaces and bounded costs
- Stochastic approximation and reinforcement learning: the interface and a little beyond
- Empirical Q-value iteration
- Decision-making problem for two-player Markov game: perspective of feedback control
- Quantifying the likelihood of learning collusive strategy equilibria
- Empirical dynamic programming
- Fitted Q-iteration by functional networks for control problems
- Asynchronous stochastic approximation with applications to average-reward reinforcement learning
- Average cost temporal-difference learning
- Deep reinforcement learning for wireless sensor scheduling in cyber-physical systems
- Analyzing anonymity attacks through noisy channels
- Optimal Distributed Uplink Channel Allocation: A Constrained MDP Formulation
- Reinforcement learning based algorithms for average cost Markov decision processes
- Dynamic pricing models for electronic business
- Stochastic approximation in non-Markovian environments
- Learning dynamic prices in electronic retail markets with customer segmentation
- Q-learning and policy iteration algorithms for stochastic shortest path problems
- Q-learning for Markov decision processes with a satisfiability criterion
- A framework for transforming specifications in reinforcement learning
- Approachability in Stackelberg stochastic games with vector costs
- Fundamental design principles for reinforcement learning algorithms
This page was built for publication: Learning algorithms for Markov decision processes with average cost
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2753225)