PulseAugur
实时 13:14:29
English(EN) Near-Optimal Stochastic Linear Bandits with Delay

新研究探索具有延迟和有界噪声的最优决策老虎机算法 · 跟踪 5 个来源

研究人员发表了关于老虎机算法的新论文,探索了在不确定性下优化决策的不同方法。一篇论文研究了具有延迟反馈的随机线性老虎机,分析了各种延迟模型如何影响遗憾保证,并将它们与多臂老虎机进行比较。另一项研究侧重于具有有界噪声的随机线性上下文老虎机,提出了一种利用集合成员估计来获得改进遗憾界限的新算法。第三篇论文研究了使用正则化稳定老虎机,推导出精确的遗憾界限和定量中心极限定理,强调了推理有效性与最优遗憾率之间的权衡。 AI

影响 这些论文推进了对老虎机算法的理论理解,有可能在面临不确定性和延迟信息的 AI 系统中实现更有效的决策。

排序理由 多篇 arXiv 论文发表了相关的机器学习研究主题。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

新研究探索具有延迟和有界噪声的最优决策老虎机算法 · 跟踪 5 个来源

报道来源 [5]

  1. arXiv cs.LG TIER_1 English(EN) · Ofir Schlisselberg, Mengxiao Zhang, Yishay Mansour ·

    近乎最优的带延迟随机线性赌博机

    arXiv:2606.16656v1 Announce Type: new Abstract: We study stochastic linear bandits with delayed feedback under several delay models and establish near-optimal regret guarantees. Our results identify when delayed linear bandits exhibit the same qualitative behavior as multi-armed …

  2. arXiv cs.LG TIER_1 English(EN) · Yishay Mansour ·

    具有延迟的近乎最优随机线性赌博机

    We study stochastic linear bandits with delayed feedback under several delay models and establish near-optimal regret guarantees. Our results identify when delayed linear bandits exhibit the same qualitative behavior as multi-armed bandits (MAB), and when the linear structure cre…

  3. arXiv stat.ML TIER_1 English(EN) · Haonan Xu, Yingying Li ·

    带噪声界限的随机线性上下文老虎机:一个集合成员方法

    arXiv:2606.20022v1 Announce Type: new Abstract: This paper considers stochastic linear contextual bandits (SLCB) with bounded reward noise. Existing works typically assume sub-Gaussian reward noise and bounded expected rewards, under which the optimal regret bound scales as $\til…

  4. arXiv stat.ML TIER_1 English(EN) · Budhaditya Halder, Ishan Sengupta, Koustav Chowdhury, Samya Praharaj, Koulik Khamaru ·

    使用正则化稳定 Bandit 算法:精确遗憾界和定量中心极限定理

    arXiv:2603.10184v2 Announce Type: replace Abstract: Statistical inference with bandit data presents fundamental challenges owing to adaptive sampling, which violates the independence assumptions underlying classical asymptotic theory. Recent work has identified stability~\citep{l…

  5. arXiv stat.ML TIER_1 English(EN) · Yingying Li ·

    带噪声界限的随机线性上下文老虎机:一种集合成员方法

    This paper considers stochastic linear contextual bandits (SLCB) with bounded reward noise. Existing works typically assume sub-Gaussian reward noise and bounded expected rewards, under which the optimal regret bound scales as $\tilde{O}(\sqrt{T})$ in terms of horizon $T$. Howeve…