An actor-critic algorithm for constrained Markov decision processes
From MaRDI portal
Recommendations
- An actor-critic algorithm with function approximation for discounted cost constrained Markov decision processes
- OnActor-Critic Algorithms
- An online actor-critic algorithm with function approximation for constrained Markov decision processes
- Actor-Critic--Type Learning Algorithms for Markov Decision Processes
- Actor-critic algorithms for hierarchical Markov decision processes
Cites work
- scientific article; zbMATH DE number 5957207 (Why is no real title available?)
- scientific article; zbMATH DE number 48727 (Why is no real title available?)
- scientific article; zbMATH DE number 1321699 (Why is no real title available?)
- scientific article; zbMATH DE number 1348599 (Why is no real title available?)
- scientific article; zbMATH DE number 700091 (Why is no real title available?)
- scientific article; zbMATH DE number 1119444 (Why is no real title available?)
- scientific article; zbMATH DE number 1786123 (Why is no real title available?)
- scientific article; zbMATH DE number 1786126 (Why is no real title available?)
- scientific article; zbMATH DE number 1786128 (Why is no real title available?)
- Actor-Critic--Type Learning Algorithms for Markov Decision Processes
- An analysis of temporal-difference learning with function approximation
- Envelope Theorems for Arbitrary Choice Sets
- Microeconomic theory
- OnActor-Critic Algorithms
- Optimal control and viscosity solutions of Hamilton-Jacobi-Bellman equations
- Stochastic approximation with two time scales
Cited in
(34)- Opportunistic Transmission over Randomly Varying Channels
- A convergent online single time scale actor critic algorithm
- A least squares temporal difference actor–critic algorithm with applications to warehouse management
- OnActor-Critic Algorithms
- Risk-Sensitive Reinforcement Learning via Policy Gradient Search
- Variance-constrained actor-critic algorithms for discounted and average reward MDPs
- Whittle index based Q-learning for restless bandits with average reward
- Actor-Critic--Type Learning Algorithms for Markov Decision Processes
- Actor-critic algorithms for hierarchical Markov decision processes
- A constrained optimization perspective on actor-critic algorithms and application to network routing
- A new learning algorithm for optimal stopping
- An Actor-Critic Algorithm With Second-Order Actor and Critic
- Dimension reduction based adaptive dynamic programming for optimal control of discrete-time nonlinear control-affine systems
- An online actor-critic algorithm with function approximation for constrained Markov decision processes
- Actor-critic algorithms based on symmetric perturbation sampling
- Quasi-Newton smoothed functional algorithms for unconstrained and constrained simulation optimization
- A note on linear function approximation using random projections
- Random search for constrained Markov decision processes with multi-policy improvement
- Multi-agent off-policy actor-critic algorithm for distributed multi-task reinforcement learning
- Finite-time analysis of natural actor-critic for POMDPs
- Stochastic approximation and reinforcement learning: the interface and a little beyond
- A primal-dual policy iteration algorithm for constrained Markov decision processes
- Artificial Intelligence and Soft Computing - ICAISC 2004
- Learning algorithms for finite horizon constrained Markov decision processes
- Optimal Distributed Uplink Channel Allocation: A Constrained MDP Formulation
- Natural actor-critic based on batch recursive least-squares
- An actor-critic algorithm with function approximation for discounted cost constrained Markov decision processes
- Risk-constrained reinforcement learning with percentile risk criteria
- Policy-based primal-dual methods for concave CMDP with variance reduction
- Finite-time analysis and restarting scheme for linear two-time-scale stochastic approximation
- A sensitivity formula for risk-sensitive cost and the actor-critic algorithm
- Approachability in Stackelberg stochastic games with vector costs
- Safety-constrained reinforcement learning with a distributional safety critic
- Delay-aware online service scheduling in high-speed railway communication systems
This page was built for publication: An actor-critic algorithm for constrained Markov decision processes
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2504518)