OnActor-Critic Algorithms
From MaRDI portal
Recommendations
Cited in
(only showing first 100 items - show all)- Reinforcement learning in the brain
- Natural actor-critic algorithms
- Estimation and approximation bounds for gradient-based reinforcement learning
- An incremental off-policy search in a model-free Markov decision process using a single sample path
- Variance-constrained actor-critic algorithms for discounted and average reward MDPs
- Totally model-free actor-critic recurrent neural-network reinforcement learning in non-Markovian domains
- Asymptotic bias of stochastic gradient search
- Reinforcement learning for a class of continuous-time input constrained optimal control problems
- Real-time reinforcement learning by sequential actor-critics and experience replay
- Convergence rate of linear two-time-scale stochastic approximation.
- Stochastic optimization for real time service capacity allocation under random service demand
- Finding intrinsic rewards by embodied evolution and constrained reinforcement learning
- Approximate stochastic annealing for online control of infinite horizon Markov decision processes
- Preference-based reinforcement learning: a formal framework and a policy iteration algorithm
- Adaptive-resolution reinforcement learning with polynomial exploration in deterministic domains
- Sell or store? An ADP approach to marketing renewable energy
- Concentration bounds for temporal difference learning with linear function approximation: the case of batch data and uniform sampling
- Fundamental design principles for reinforcement learning algorithms
- Mixed density methods for approximate dynamic programming
- Multi-agent reinforcement learning: a selective overview of theories and algorithms
- Neural circuits for learning context-dependent associations of stimuli
- Weak convergence of dynamical systems in two timescales
- Control strategy of speed servo systems based on deep reinforcement learning
- TD-regularized actor-critic methods
- Performance optimization for a class of generalized stochastic Petri nets
- Reinforcement learning for a biped robot based on a CPG-actor-critic method
- A tutorial on the cross-entropy method
- Linear stochastic approximation driven by slowly varying Markov chains
- An actor-critic algorithm for constrained Markov decision processes
- Dynamic programming and suboptimal control: a survey from ADP to MPC
- Reinforcement learning based algorithms for average cost Markov decision processes
- Two-timescale stochastic gradient descent in continuous time with applications to joint online parameter estimation and optimal sensor placement
- Hebbian versus gradient training of ESN actors in closed-loop ACD
- A constrained optimization perspective on actor-critic algorithms and application to network routing
- A convergent online single time scale actor critic algorithm
- Actor-critic algorithms based on symmetric perturbation sampling
- Multiscale Q-learning with linear function approximation
- Efficient model-based reinforcement learning for approximate online optimal control
- Natural actor-critic based on batch recursive least-squares
- A Spiking Neural Network Model of an Actor-Critic Learning Agent
- Dynamic treatment regimes: technical challenges and applications
- An online actor-critic algorithm with function approximation for constrained Markov decision processes
- Stabilization of stochastic approximation by step size adaptation
- Risk-constrained reinforcement learning with percentile risk criteria
- From infinite to finite programs: explicit error bounds with applications to approximate dynamic programming
- Autonomous reinforcement learning with experience replay
- Artificial Intelligence and Soft Computing - ICAISC 2004
- Actor-Critic--Type Learning Algorithms for Markov Decision Processes
- Finite-time analysis and restarting scheme for linear two-time-scale stochastic approximation
- Actor-critic method for high dimensional static Hamilton-Jacobi-Bellman partial differential equations based on neural networks
- scientific article; zbMATH DE number 7625165 (Why is no real title available?)
- A new approximate dynamic programming algorithm based on an actor–critic framework for optimal control of alkali–surfactant–polymer flooding
- Actor-Critic–Like Stochastic Adaptive Search for Continuous Simulation Optimization
- What is the value of the cross-sectional approach to deep reinforcement learning?
- Simple and optimal methods for stochastic variational inequalities. II: Markovian noise and policy evaluation in reinforcement learning
- Queueing network controls via deep reinforcement learning
- Global convergence of policy gradient methods to (almost) locally optimal policies
- Deep Reinforcement Learning: A State-of-the-Art Walkthrough
- Policy optimization for \(\mathcal{H}_2\) linear control with \(\mathcal{H}_\infty\) robustness guarantee: implicit regularization and global convergence
- Asynchronous stochastic approximation with differential inclusions
- On the convergence of simulation-based iterative methods for solving singular linear systems
- Derivatives of logarithmic stationary distributions for policy gradient reinforcement learning
- Two Time-Scale Stochastic Approximation with Controlled Markov Noise and Off-Policy Temporal-Difference Learning
- A perturbation approach to approximate value iteration for average cost Markov decision processes with Borel spaces and bounded costs.
- Approximation of average cost Markov decision processes using empirical distributions and concentration inequalities
- Actor-critic algorithms with online feature adaptation
- An Actor-Critic Algorithm With Second-Order Actor and Critic
- Reinforcement Learning, Spike-Time-Dependent Plasticity, and the BCM Rule
- Asymptotic analysis of temporal-difference learning algorithms with constant step-sizes
- Asymptotic analysis of temporal-difference learning algorithms with constant step-sizes
- A sensitivity formula for risk-sensitive cost and the actor-critic algorithm
- A Small Gain Analysis of Single Timescale Actor Critic
- Toward multi-target self-organizing pursuit in a partially observable Markov game
- Approximate Newton Policy Gradient Algorithms
- An Improved Unconstrained Approach for Bilevel Optimization
- Bayesian sequential optimal experimental design for nonlinear models using policy gradient reinforcement learning
- Non-iterative generation of an optimal mesh for a blade passage using deep reinforcement learning
- An actor-critic algorithm with policy gradients to solve the job shop scheduling problem using deep double recurrent agents
- Reward-respecting subtasks for model-based reinforcement learning
- Multi-agent off-policy actor-critic algorithm for distributed multi-task reinforcement learning
- Smoothing policies and safe policy gradients
- Variational actor-critic algorithms,
- Softmax policy gradient methods can take exponential time to converge
- On the sample complexity of actor-critic method for reinforcement learning with function approximation
- Geometry and convergence of natural policy gradient methods
- Tutorial on Amortized Optimization
- Recent advances in reinforcement learning in finance
- Multi-agent natural actor-critic reinforcement learning algorithms
- An accelerated proximal algorithm for regularized nonconvex and nonsmooth bi-level optimization
- An actor-critic algorithm with function approximation for discounted cost constrained Markov decision processes
- Error controlled actor-critic
- On centralized critics in multi-agent reinforcement learning
- Actor prioritized experience replay
- Global convergence of natural policy gradient with Hessian-aided momentum variance reduction
- Finite-time analysis of natural actor-critic for POMDPs
- A stabilizing reinforcement learning approach for sampled systems with partially unknown models
- Stochastic approximation and reinforcement learning: the interface and a little beyond
- Blackbox simulation optimization
- Exploring reinforcement learning in process control: a comprehensive survey
- Deep reinforcement learning for infinite horizon mean field problems in continuous spaces
This page was built for publication: OnActor-Critic Algorithms
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4443033)