The apparent conflict between estimation and control - a survey of the two-armed bandit problem
From MaRDI portal
Publication:1226072
Cites work
- A note on the two-armed bandit problem with finite memory
- A SEQUENTIAL DECISION PROBLEM WITH A FINITE MEMORY
- A Sequential Design for the Two Armed Bandit
- A study of the relationship between identification and optimization in adaptive control problems
- Contributions to the "Two-Armed Bandit" Problem
- Dual control theory. II
- Finite-memory hypothesis testing--A critique (Corresp.)
- Finite-memory hypothesis testing--Comments on a critique (Corresp.)
- Finite-Time Performance of Some Two-Armed Bandit Controllers
- scientific article; zbMATH DE number 3125136 (Why is no real title available?)
- scientific article; zbMATH DE number 3152611 (Why is no real title available?)
- scientific article; zbMATH DE number 3168214 (Why is no real title available?)
- scientific article; zbMATH DE number 3045589 (Why is no real title available?)
- scientific article; zbMATH DE number 3059214 (Why is no real title available?)
- scientific article; zbMATH DE number 3109895 (Why is no real title available?)
- Hypothesis testing with finite memory in finite time (Corresp.)
- Hypothesis Testing with Finite Statistics
- Learning Automata - A Survey
- Learning with Finite Memory
- On a Problem of Robbins
- On Memory Saved by Randomization
- On Sequential Designs for Maximizing the Sum of n Observations
- On the Asymptotic Performances of Finite-State Two-Armed Bandit Controllers
- On the Theory of Apportionment
- Randomized Rules for the Two-Armed-Bandit with Finite Memory
- Reply to 'Finite memory hypothesis testing - Comments on a critique' by Cover, T.M., and Hellman, M.E.
- Some aspects of the sequential design of experiments
- Testing a simple symmetric hypothesis by a finite-memory deterministic algorithm
- The effects of randomization on finite-memory decision schemes
- The Robbins-Isbell Two-Armed-Bandit Problem with Finite Memory
- The two-armed-bandit problem with time-invariant finite memory
Cited in
(6)- The N-armed bandit with unimodal structure
- On a general class of absorbing-barrier learning algorithms
- epsilon-optimality of a general class of learning algorithms
- Reliability of internal prediction/estimation and its application. I: Adaptive action selection reflecting reliability of value function
- Opportunistic spectrum access in unslotted primary systems
- Multiple objective optimization approach to adaptive and learning control
This page was built for publication: The apparent conflict between estimation and control - a survey of the two-armed bandit problem
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q1226072)