Learning action probabilities from delayed reinforcement
From MaRDI portal
Recommendations
- A new approach to the design of reinforcement schemes for learning automata
- A tutorial survey of reinforcement learning
- The asymptotic optimality of discretized linear reward-inaction learning automata
- Absorbing and Ergodic Discretized Two-Action Learning Automata
- Learning automata processing ergodicity of the mean: The two-action case
Cited in
(7)- Learning with incomplete information and the mathematical structure behind it
- A special learning process with time delay
- A new approach to the design of reinforcement schemes for learning automata
- scientific article; zbMATH DE number 67800 (Why is no real title available?)
- scientific article; zbMATH DE number 1479845 (Why is no real title available?)
- scientific article; zbMATH DE number 1356140 (Why is no real title available?)
- scientific article; zbMATH DE number 234518 (Why is no real title available?)
This page was built for publication: Learning action probabilities from delayed reinforcement
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4278272)