Dynamic learning and decision making via basis weight vectors
From MaRDI portal
Recommendations
- Contextual bandits with continuous actions: smoothing, zooming, and adapting
- Optimal timing of decisions: a general theory based on continuation values
- The Continuum-Armed Bandit Problem
- Online learning in Markov decision processes with continuous actions
- Statistical inference for online decision making: in a contextual bandit setting
Cites work
- 10.1162/153244303321897663
- A Learning Approach for Interactive Marketing to a Customer Segment
- A partially observed Markov decision process for dynamic pricing
- A survey of algorithmic methods for partially observed Markov decision processes
- A survey of solution techniques for the partially observed Markov decision process
- Approximate dynamic programming. Solving the curses of dimensionality
- Controlling a Stochastic Process with Unknown Parameters
- Dynamic assortment with demand learning for seasonal consumer goods
- Dynamic pricing for nonperishable products with demand learning
- Dynamic pricing under a general parametric choice model
- Dynamic pricing with a prior on market response
- Dynamic pricing without knowing the demand function: risk bounds and near-optimal algorithms
- Dynamic pricing: a learning approach
- Dynamic programming and optimal control. Vol. 1.
- Dynamic programming and optimal control. Vol. 2
- Dynamic selling mechanisms for product differentiation and learning
- scientific article; zbMATH DE number 3638998 (Why is no real title available?)
- scientific article; zbMATH DE number 700091 (Why is no real title available?)
- scientific article; zbMATH DE number 775283 (Why is no real title available?)
- Implementation and parallelization of a reverse-search algorithm for Minkowski sums
- Information relaxations and duality in stochastic dynamic programs
- Information Relaxations, Duality, and Convex Stochastic Dynamic Programs
- Investment timing with incomplete information and multiple means of learning
- Linearly parameterized bandits
- On incomplete learning and certainty-equivalence control
- Optimal Experimentation in a Changing Environment
- Partially observable Markov decision processes: a geometric technique and analysis
- Partially observed Markov decision processes. From filtering to controlled sensing
- Reinforcement learning. An introduction
- Sequential Tests of Statistical Hypotheses
- Some aspects of the sequential design of experiments
- State of the Art—A Survey of Partially Observable Markov Decision Processes: Theory, Models, and Algorithms
- The Optimal Control of Partially Observable Markov Processes over a Finite Horizon
Cited in
(6)- Optimal timing of decisions: a general theory based on continuation values
- Dynamic benchmark targeting
- Dynamic parameters in sequential decision making
- Multidimensional binary search for contextual decision-making
- Contextual search via intrinsic volumes
- Statistical inference for online decision making: in a contextual bandit setting
This page was built for publication: Dynamic learning and decision making via basis weight vectors
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5095179)