PulseAugur
实时 08:56:59
English(EN) Sequential Batch Learning in Finite-Action Linear Contextual Bandits

arXiv论文通过记忆和批量学习改进上下文老虎机算法

两篇arXiv论文深入探讨了上下文老虎机算法的复杂性。第一篇论文《具有记忆的广义线性老虎机》改进了线性和广义线性模型的遗憾界限,尽管存在非线性奖励和记忆效应,仍达到了 $\tilde{O}(\sqrt{T})$ 的速率。第二篇论文《有限动作线性上下文老虎机中的顺序批量学习》解决了批量决策的挑战,为具有固定批量约束的场景建立了遗憾界限和算法,适用于个性化治疗和推荐系统等领域。 AI

影响 这些论文在不确定性下的序贯决策方面推进了理论理解和算法方法,可能改进个性化系统。

排序理由 两篇发表在arXiv上的学术论文,详细介绍了上下文老虎机算法的进展。

在 arXiv stat.ML 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

arXiv论文通过记忆和批量学习改进上下文老虎机算法

报道来源 [4]

  1. arXiv cs.LG TIER_1 English(EN) · Yuhao Liu, Yu Chen, Longbo Huang ·

    具有任意延迟的Lipschitz Bandit

    arXiv:2608.15036v1 Announce Type: new Abstract: The Lipschitz bandit problem extends the traditional multi-armed bandit framework to continuous action spaces by assuming that the reward functions satisfy a Lipschitz condition. This work investigates Lipschitz bandits under arbitr…

  2. arXiv cs.LG TIER_1 English(EN) · Anling Xiang, Yuwen Yang, Yang Shen ·

    超越峰值积压:容量受限延迟老虎机优化中的条件能量与时间几何

    arXiv:2608.16216v1 Announce Type: new Abstract: What is the right delay complexity when a learner can track only $C$ pending feedback items and discarded feedback is permanently lost? Existing one-point bandit convex optimization guarantees in this model pay $\sqrt{T\sigma_{\max}…

  3. arXiv stat.ML TIER_1 English(EN) · Heesang Ann, Hyunjun Choi, Taehyun Hwang, Younghoon Shin, Haeju Cheong, Min-hwan Oh ·

    具有记忆的广义线性老虎机

    arXiv:2608.15848v1 Announce Type: new Abstract: We study generalized linear bandits with memory, an endogenous non-stationary setting in which rewards depend on past actions through a finite memory matrix. Building on prior work for linear models (Clerici et al., 2024), we show t…

  4. arXiv stat.ML TIER_1 English(EN) · Yanjun Han, Zhengqing Zhou, Zihao Hu, Jose Blanchet, Peter W. Glynn, Yinyu Ye, Zhengyuan Zhou ·

    有限动作线性上下文老虎机中的顺序批量学习

    arXiv:2004.06321v2 Announce Type: replace-cross Abstract: We study the sequential batch learning problem in linear contextual bandits with finite action sets, where the decision maker is constrained to split incoming individuals into (at most) a fixed number of batches and can on…