A two-step algorithm for learning from unspecific reinforcement
From MaRDI portal
Abstract: We study a simple learning model based on the Hebb rule to cope with "delayed", unspecific reinforcement. In spite of the unspecific nature of the information-feedback, convergence to asymptotically perfect generalization is observed, with a rate depending, however, in a non- universal way on learning parameters. Asymptotic convergence can be as fast as that of Hebbian learning, but may be slower. Moreover, for a certain range of parameter settings, it depends on initial conditions whether the system can reach the regime of asymptotically perfect generalization, or rather approaches a stationary state of poor generalization.
Recommendations
- Learning structured data from unspecific reinforcement
- Learning with incomplete information and the mathematical structure behind it
- Simple statistical gradient-following algorithms for connectionist reinforcement learning
- Hebbian errors in learning: an analysis using the Oja model
- Combining Hebbian and reinforcement learning in a minibrain model
Cited in
(2)
This page was built for publication: A two-step algorithm for learning from unspecific reinforcement
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4947717)