PulseAugur
实时 10:59:19
English(EN) Sequential Batch Learning in Finite-Action Linear Contextual Bandits

新研究通过延迟和内存解决复杂的老虎机问题 · 已追踪 4 个来源

arXiv 上发表了四篇新的研究论文,探讨了老虎机算法的进展,重点关注反馈延迟、容量限制和内存集成等挑战。这些研究引入了新颖的算法和分析,以改进各种设置下的遗憾界限,包括具有任意延迟的 Lipschitz 老虎机、容量受限优化以及具有内存的广义线性老虎机。其中一篇论文专门讨论了线性上下文老虎机中的顺序批量学习,为个性化决策问题提供了更细粒度的表述。 AI

影响 这些论文在不确定性下的复杂决策问题的理论理解和算法解决方案方面取得了进展,可能影响需要延迟或有限反馈的高效学习的领域。

排序理由 arXiv 上发表了多篇学术论文,详细介绍了老虎机问题的新算法和分析。

在 arXiv stat.ML 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

新研究通过延迟和内存解决复杂的老虎机问题 · 已追踪 4 个来源

报道来源 [4]

  1. arXiv cs.LG TIER_1 English(EN) · Yuhao Liu, Yu Chen, Longbo Huang ·

    具有任意延迟的Lipschitz Bandit

    arXiv:2608.15036v1 Announce Type: new Abstract: The Lipschitz bandit problem extends the traditional multi-armed bandit framework to continuous action spaces by assuming that the reward functions satisfy a Lipschitz condition. This work investigates Lipschitz bandits under arbitr…

  2. arXiv cs.LG TIER_1 English(EN) · Anling Xiang, Yuwen Yang, Yang Shen ·

    超越峰值积压:容量受限延迟老虎机优化中的条件能量与时间几何

    arXiv:2608.16216v1 Announce Type: new Abstract: What is the right delay complexity when a learner can track only $C$ pending feedback items and discarded feedback is permanently lost? Existing one-point bandit convex optimization guarantees in this model pay $\sqrt{T\sigma_{\max}…

  3. arXiv stat.ML TIER_1 English(EN) · Heesang Ann, Hyunjun Choi, Taehyun Hwang, Younghoon Shin, Haeju Cheong, Min-hwan Oh ·

    具有记忆的广义线性老虎机

    arXiv:2608.15848v1 Announce Type: new Abstract: We study generalized linear bandits with memory, an endogenous non-stationary setting in which rewards depend on past actions through a finite memory matrix. Building on prior work for linear models (Clerici et al., 2024), we show t…

  4. arXiv stat.ML TIER_1 English(EN) · Yanjun Han, Zhengqing Zhou, Zihao Hu, Jose Blanchet, Peter W. Glynn, Yinyu Ye, Zhengyuan Zhou ·

    有限动作线性上下文老虎机中的顺序批量学习

    arXiv:2004.06321v2 Announce Type: replace-cross Abstract: We study the sequential batch learning problem in linear contextual bandits with finite action sets, where the decision maker is constrained to split incoming individuals into (at most) a fixed number of batches and can on…