Reinforcement with fading memories
From MaRDI portal
Abstract: We study the effect of imperfect memory on decision making in the context of a stochastic sequential action-reward problem. An agent chooses a sequence of actions which generate discrete rewards at different rates. She is allowed to make new choices at rate , while past rewards disappear from her memory at rate . We focus on a family of decision rules where the agent makes a new choice by randomly selecting an action with a probability approximately proportional to the amount of past rewards associated with each action in her memory. We provide closed-form formulae for the agent's steady-state choice distribution in the regime where the memory span is large (), and show that the agent's success critically depends on how quickly she updates her choices relative to the speed of memory decay. If , the agent almost always chooses the best action, i.e., the one with the highest reward rate. Conversely, if , the agent chooses an action with a probability roughly proportional to its reward rate.
Recommendations
Cites work
- A Markov model for the spread of viruses in an open population
- A survey of random processes with reinforcement
- An ODE for an overloaded X model involving a stochastic averaging principle
- Chasing demand: learning and earning in a changing environment
- Differential equation approximations for Markov chains
- Discrete Choice Methods with Simulation
- scientific article; zbMATH DE number 3152611 (Why is no real title available?)
- scientific article; zbMATH DE number 1782874 (Why is no real title available?)
- scientific article; zbMATH DE number 847278 (Why is no real title available?)
- scientific article; zbMATH DE number 2237386 (Why is no real title available?)
- scientific article; zbMATH DE number 3109695 (Why is no real title available?)
- Learning in games via reinforcement and regularization
- Learning with Finite Memory
- Non-stationary stochastic optimization
- On the capacity of information processing systems
- On the convergence of reinforcement learning
- On the power of (even a little) resource pooling
- Optimal properties of stimulus-response learning models.
- Optional sampling of submartingales indexed by partially ordered sets
- Solutions of ordinary differential equations as limits of pure jump markov processes
- State space collapse with application to heavy traffic limits for multiclass queueing networks
- Strong approximation for Markovian service networks
- Strong approximation theorems for density dependent Markov chains
- The Nonstochastic Multiarmed Bandit Problem
- The two-armed-bandit problem with time-invariant finite memory
- To queue or not to queue: equilibrium behavior in queueing systems.
- Uniform acceleration expansions for Markov chains with time-varying rates
This page was built for publication: Reinforcement with fading memories
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3387923)