Simple statistical gradient-following algorithms for connectionist reinforcement learning
From MaRDI portal
Recommendations
Cites work
- A new approach to the design of reinforcement schemes for learning automata
- An N-player sequential stochastic game with identical payoffs
- Associative search network: A reinforcement learning associative memory
- Decentralized learning in finite Markov chains
- scientific article; zbMATH DE number 4066707 (Why is no real title available?)
- scientific article; zbMATH DE number 3657150 (Why is no real title available?)
- scientific article; zbMATH DE number 3551675 (Why is no real title available?)
- Pattern-recognizing stochastic learning automata
Cited in
(only showing first 100 items - show all)- Reinforcement learning in the brain
- Natural actor-critic algorithms
- Autonomous vehicle navigation using evolutionary reinforcement learning
- Estimation and approximation bounds for gradient-based reinforcement learning
- Variance-constrained actor-critic algorithms for discounted and average reward MDPs
- Optimal node perturbation in linear perceptrons with uncertain eligibility trace
- Reliability of internal prediction/estimation and its application. I: Adaptive action selection reflecting reliability of value function
- Continuous action set learning automata for stochastic optimization
- Two forms of immediate reward reinforcement learning for exploratory data analysis
- Preference-based reinforcement learning: a formal framework and a policy iteration algorithm
- A projected primal-dual gradient optimal control method for deep reinforcement learning
- Novelty detection improves performance of reinforcement learners in fluctuating, partially observable environments
- Importance sampling in reinforcement learning with an estimated behavior policy
- HNS: hierarchical negative sampling for network representation learning
- Revisiting the ODE method for recursive algorithms: fast convergence using quasi stochastic approximation
- Dealing with multiple experts and non-stationarity in inverse reinforcement learning: an application to real-life problems
- Deep reinforcement learning for inventory control: a roadmap
- Risk-averse policy optimization via risk-neutral policy optimization
- Measurement error models: from nonparametric methods to deep neural networks
- Neural large neighborhood search for routing problems
- Multi-agent reinforcement learning: a selective overview of theories and algorithms
- Opportunities for reinforcement learning in stochastic dynamic vehicle routing
- Heavy-tails and randomized restarting beam search in goal-oriented neural sequence decoding
- Semi-discrete optimization through semi-discrete optimal transport: a framework for neural architecture search
- Learning the travelling salesperson problem requires rethinking generalization
- Learning to compute the metric dimension of graphs
- Hybrid offline/online optimization for energy management via reinforcement learning
- Approximate Bayesian model inversion for PDEs with heterogeneous and state-dependent coefficients
- Deep reinforcement learning for the optimal placement of cryptocurrency limit orders
- Enhance load forecastability: optimize data sampling policy by reinforcing user behaviors
- Smoothed functional-based gradient algorithms for off-policy reinforcement learning: a non-asymptotic viewpoint
- A review on deep reinforcement learning for fluid mechanics
- Model-based reinforcement learning with dimension reduction
- Branes with brains: exploring string vacua with deep reinforcement learning
- Compatible natural gradient policy search
- TD-regularized actor-critic methods
- Learning to attend: modeling the shaping of selectivity in infero-temporal cortex in a categorization task
- Learning flexible sensori-motor mappings in a complex network
- Reinforcement learning for a biped robot based on a CPG-actor-critic method
- Restricted gradient-descent algorithm for value-function approximation in reinforcement learning
- A study of mechanisms for improving robotic group performance
- Model-based contextual policy search for data-efficient generalization of robot skills
- Constructing effective personalized policies using counterfactual inference from biased data sets with many features
- Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithm
- Varieties of Helmholtz machine
- A reinforcement learning approach to the orienteering problem with time windows
- Reinforcement learning for combinatorial optimization: a survey
- Rationalizing predictions by adversarial information calibration
- Dynamic graph conv-LSTM model with dynamic positional encoding for the large-scale traveling salesman problem
- Reinforcement learning theory, algorithms and its application
- scientific article; zbMATH DE number 1708090 (Why is no real title available?)
- Adaptive learning via selectionism and Bayesianism. I: Connection between the two
- Adaptive playouts for online learning of policies during Monte Carlo tree search
- Using Expectation-Maximization for Reinforcement Learning
- Environment-driven distributed evolutionary adaptation in a population of autonomous robotic agents
- A stochastic policy search model for matching behavior
- Active inference and agency: optimal control without cost functions
- Posterior weighted reinforcement learning with state uncertainty
- Recurrent policy gradients
- scientific article; zbMATH DE number 67800 (Why is no real title available?)
- Policy search for motor primitives in robotics
- scientific article; zbMATH DE number 1966632 (Why is no real title available?)
- Analysis and improvement of policy gradient estimation
- scientific article; zbMATH DE number 6982909 (Why is no real title available?)
- Risk-constrained reinforcement learning with percentile risk criteria
- GSNs: generative stochastic networks
- Autonomous reinforcement learning with experience replay
- Neural architecture search: a survey
- Non-parametric policy search with limited information loss
- A SELF-IMPROVING FUZZY CEREBELLAR MODEL ARTICULATION CONTROLLER WITH STOCHASTIC ACTION GENERATION
- Stochastic dynamics of reinforcement learning
- Learning automata in feedforward connectionist systems
- scientific article; zbMATH DE number 1424385 (Why is no real title available?)
- A two-step algorithm for learning from unspecific reinforcement
- Greedy attack and Gumbel attack: generating adversarial examples for discrete data
- Ancestral Gumbel-top-k sampling for sampling without replacement
- \textsc{NeVAE}: a deep generative model for molecular graphs
- scientific article; zbMATH DE number 7370547 (Why is no real title available?)
- scientific article; zbMATH DE number 7370594 (Why is no real title available?)
- On the theory of policy gradient methods: optimality, approximation, and distribution shift
- Reinforcement learning in sparse-reward environments with hindsight policy gradients
- A human-centered data-driven planner-actor-critic architecture via logic programming
- High generalization performance structured self-attention model for knapsack problem
- Task-aware verifiable RNN-based policies for partially observable Markov decision processes
- scientific article; zbMATH DE number 7625182 (Why is no real title available?)
- Actor-Critic–Like Stochastic Adaptive Search for Continuous Simulation Optimization
- Bayesian Variational Inference for Exponential Random Graph Models
- Stochastic learning approach for binary optimization: application to Bayesian optimal design of experiments
- Supervised Visual Attention for Simultaneous Multimodal Machine Translation
- Fast global convergence of natural policy gradient methods with entropy regularization
- A novel online gait optimization approach for biped robots with point-feet
- Global convergence of policy gradient methods to (almost) locally optimal policies
- Deep Reinforcement Learning: A State-of-the-Art Walkthrough
- Robust reinforcement learning with Bayesian optimisation and quadrature
- scientific article; zbMATH DE number 7307467 (Why is no real title available?)
- Full gradient DQN reinforcement learning: a provably convergent scheme
- Set-to-Sequence Methods in Machine Learning: A Review
- Dynamic neural Turing machine with continuous and discrete addressing schemes
- A learning framework for winner-take-all networks with stochastic synapses
- Adaptive learning algorithm convergence in passive and reactive environments
This page was built for publication: Simple statistical gradient-following algorithms for connectionist reinforcement learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q1812928)