An actor-critic algorithm for constrained Markov decision processes
From MaRDI portal
Recommendations
- An actor-critic algorithm with function approximation for discounted cost constrained Markov decision processes
- OnActor-Critic Algorithms
- An online actor-critic algorithm with function approximation for constrained Markov decision processes
- Actor-Critic--Type Learning Algorithms for Markov Decision Processes
- Actor-critic algorithms for hierarchical Markov decision processes
Cites work
- Actor-Critic--Type Learning Algorithms for Markov Decision Processes
- An analysis of temporal-difference learning with function approximation
- Envelope Theorems for Arbitrary Choice Sets
- scientific article; zbMATH DE number 5957207 (Why is no real title available?)
- scientific article; zbMATH DE number 48727 (Why is no real title available?)
- scientific article; zbMATH DE number 1321699 (Why is no real title available?)
- scientific article; zbMATH DE number 1348599 (Why is no real title available?)
- scientific article; zbMATH DE number 700091 (Why is no real title available?)
- scientific article; zbMATH DE number 1119444 (Why is no real title available?)
- scientific article; zbMATH DE number 1786123 (Why is no real title available?)
- scientific article; zbMATH DE number 1786126 (Why is no real title available?)
- scientific article; zbMATH DE number 1786128 (Why is no real title available?)
- Microeconomic theory
- OnActor-Critic Algorithms
- Optimal control and viscosity solutions of Hamilton-Jacobi-Bellman equations
- Stochastic approximation with two time scales
Cited in
(36)- Variance-constrained actor-critic algorithms for discounted and average reward MDPs
- Approachability in Stackelberg stochastic games with vector costs
- Delay-aware online service scheduling in high-speed railway communication systems
- Whittle index based Q-learning for restless bandits with average reward
- Learning algorithms for finite horizon constrained Markov decision processes
- A note on linear function approximation using random projections
- A constrained optimization perspective on actor-critic algorithms and application to network routing
- A convergent online single time scale actor critic algorithm
- Actor-critic algorithms based on symmetric perturbation sampling
- A least squares temporal difference actor–critic algorithm with applications to warehouse management
- Natural actor-critic based on batch recursive least-squares
- Opportunistic Transmission over Randomly Varying Channels
- An online actor-critic algorithm with function approximation for constrained Markov decision processes
- OnActor-Critic Algorithms
- Risk-constrained reinforcement learning with percentile risk criteria
- Artificial Intelligence and Soft Computing - ICAISC 2004
- Actor-Critic--Type Learning Algorithms for Markov Decision Processes
- Finite-time analysis and restarting scheme for linear two-time-scale stochastic approximation
- Risk-Sensitive Reinforcement Learning via Policy Gradient Search
- Optimal Distributed Uplink Channel Allocation: A Constrained MDP Formulation
- Quasi-Newton smoothed functional algorithms for unconstrained and constrained simulation optimization
- An Actor-Critic Algorithm With Second-Order Actor and Critic
- A sensitivity formula for risk-sensitive cost and the actor-critic algorithm
- Dimension reduction based adaptive dynamic programming for optimal control of discrete-time nonlinear control-affine systems
- Multi-agent off-policy actor-critic algorithm for distributed multi-task reinforcement learning
- Safety-constrained reinforcement learning with a distributional safety critic
- An actor-critic algorithm with function approximation for discounted cost constrained Markov decision processes
- Finite-time analysis of natural actor-critic for POMDPs
- Stochastic approximation and reinforcement learning: the interface and a little beyond
- A primal-dual policy iteration algorithm for constrained Markov decision processes
- Policy-based primal-dual methods for concave CMDP with variance reduction
- An online value iteration method for stochastic linear quadratic control with multiplicative noise
- Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
- A new learning algorithm for optimal stopping
- Actor-critic algorithms for hierarchical Markov decision processes
- Random search for constrained Markov decision processes with multi-policy improvement
This page was built for publication: An actor-critic algorithm for constrained Markov decision processes
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2504518)