PulseAugur
实时 21:11:05
English(EN) Thompson Sampling Is 2-Competitive for Mistakes

Thompson Sampling 在贝叶斯老虎机模型中被证明在错误方面具有 2-竞争性

一篇新发表在 arXiv 上的论文详细介绍了一个在贝叶斯老虎机模型方面的理论进展,证明了 Thompson sampling 在错误方面具有 2-竞争性。这意味着 Thompson sampling 所犯的错误预期数量最多是任何其他策略的两倍。该分析适用于独立的潜在臂(arm)过程,其中臂仅在被抽取时才会演变,证实了 GuhaMunagala 在 2014 年对随机老虎机提出的猜想。该结果适用于各种加权方案,包括固定视野和几何贴现。 AI

影响 为强化学习和不确定性下的决策制定中的常用算法提供了理论保证。

排序理由 学术论文,详细介绍了机器学习方面的理论结果。

在 arXiv stat.ML 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

Thompson Sampling 在贝叶斯老虎机模型中被证明在错误方面具有 2-竞争性

报道来源 [2]

  1. arXiv stat.ML TIER_1 English(EN) · Mark Sellke, Gregory Valiant ·

    Thompson Sampling Is 2-Competitive for Mistakes

    arXiv:2607.12389v1 Announce Type: new Abstract: We consider Bayesian bandit models and prove that Thompson sampling makes at most twice the expected number of mistakes (selections of a suboptimal arm) as any other policy. Our analysis applies as long as the latent arm processes a…

  2. arXiv stat.ML TIER_1 English(EN) · Gregory Valiant ·

    Thompson Sampling Is 2-Competitive for Mistakes

    We consider Bayesian bandit models and prove that Thompson sampling makes at most twice the expected number of mistakes (selections of a suboptimal arm) as any other policy. Our analysis applies as long as the latent arm processes are independent and each arm evolves only when pl…