本研究论文调查了 1/2-Tsallis-INF 算法(一种用于最小化多臂老虎机问题遗憾值的已知方法)在应用于识别最佳臂任务时的有效性。该研究分析了该算法在随机老虎机设置下的失效概率,重点关注次优臂的采样方式以及这对累积损失估计器的影响。通过为估计累积损失的间隙过程开发一个李雅普诺夫函数,论文为失效概率建立了多项式上限,并证明了这些上限中的指数基本上是紧密的。 AI
排序理由 该条目是一篇学术论文,详细介绍了算法的理论分析。[lever_c_demoted from research: ic=1 ai=1.0]
- 1/2-Tsallis-INF
- Adversarial Bandits With Multi-User Delayed Feedback: Theory and Application
- alphaXiv
- Best arm identification
- CatalyzeX
- DagsHub
- Diffusion toy model
- Follow-the-Regularized-Leader
- Gotit.pub
- Hugging Face
- Lyapunov function
- Multi-armed bandits for adjudicating documents in pooling-based evaluation of information retrieval systems
- Regret-minimization algorithms for multi-agent cooperative learning systems
- ScienceCast
- Stochastic bandits with arm-dependent delays
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →