This paper introduces a novel stochastic stability framework to analyze Thompson Sampling (TS) algorithms in dynamic decision-making scenarios where the underlying model might be misspecified. The research provides a detailed classification of posterior evolution in a two-armed Gaussian bandit, identifying distinct regimes that predict limiting beliefs, action frequencies, and asymptotic regret. The framework is then generalized to finite model classes, offering a qualitative and geometric understanding of TS behavior under misspecification and laying groundwork for robust decision-making in structured bandit problems. AI
IMPACT Provides a theoretical foundation for robust decision-making in machine learning algorithms facing uncertain environments.
RANK_REASON The item is an academic paper published on arXiv detailing a new theoretical framework and analysis of an existing algorithm. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- belief simplex
- CatalyzeX
- DagsHub
- Daniel Chen
- Gaussian bandit
- Gotit.pub
- Hugging Face
- Influence Flower
- Markov chain
- ScienceCast
- Thompson sampling
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →