Robust experimentation in the continuous time bandit problem
From MaRDI portal
Abstract: We study the experimentation dynamics of a decision maker (DM) in a two-armed bandit setup (Bolton and Harris (1999)), where the agent holds ambiguous beliefs regarding the distribution of the return process of one arm and is certain about the other one. The DM entertains Multiplier preferences a la Hansen and Sargent (2001), thus we frame the decision making environment as a two-player differential game against nature in continuous time. We characterize the DM value function and her optimal experimentation strategy that turns out to follow a cut-off rule with respect to her belief process. The belief threshold for exploring the ambiguous arm is found in closed form and is shown to be increasing with respect to the ambiguity aversion index. We then study the effect of provision of an unambiguous information source about the ambiguous arm. Interestingly, we show that the exploration threshold rises unambiguously as a result of this new information source, thereby leading to more conservatism. This analysis also sheds light on the efficient time to reach for an expert opinion.
Recommendations
Cites work
- scientific article; zbMATH DE number 3638998 (Why is no real title available?)
- scientific article; zbMATH DE number 2235418 (Why is no real title available?)
- A Corrected Proof of the Stochastic Verification Theorem within the Framework of Viscosity Solutions
- Ambiguity Aversion, Robustness, and the Variational Representation of Preferences
- Ambiguity aversion in multi-armed bandit problems
- Ambiguity sharing and the lack of relative performance evaluation
- Dynamic variational preferences
- Erratum: ``A corrected proof of the stochastic verification theorem within the framework of viscosity solutions
- Learning Under Ambiguity
- Learning from ambiguous urns
- Learning to disagree in a game of experimentation
- Maxmin expected utility with non-unique prior
- Optimal Experimentation in a Changing Environment
- Optimal Search for the Best Alternative
- Optimal Stopping With Multiple Priors
- Optimal control of diffusion processes and hamilton–jacobi–bellman equations part 2 : viscosity solutions and uniqueness
- Optimal stopping under ambiguity in continuous time
- Recursive multiple-priors.
- Robust Contracts in Continuous Time
- Robust control and model misspecification
- Robustness and ambiguity in continuous time
- Sequential Choice Under Ambiguity: Intuitive Solutions to the Armed-Bandit Problem
- Some Properties of Viscosity Solutions of Hamilton-Jacobi Equations
- Stochastic Verification Theorems within the Framework of Viscosity Solutions
- Strategic Experimentation
- Strategic Experimentation with Exponential Bandits
- Strategic experimentation with private payoffs
- The K-armed bandit problem with multiple priors
Cited in
(8)- Robustness of stochastic bandit policies
- Dual auctions for assigning winners and compensating losers
- Exploration and correlation
- Optimal learning under robustness and time-consistency
- The K-armed bandit problem with multiple priors
- Reputation, learning and project choice in frictional economies
- Ambiguity aversion in multi-armed bandit problems
- A note on optimal experimentation under risk aversion
This page was built for publication: Robust experimentation in the continuous time bandit problem
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2150441)