English(EN)Sequential Batch Learning in Finite-Action Linear Contextual Bandits
新研究通过延迟和内存解决复杂的老虎机问题 · 已追踪 4 个来源
作者PulseAugur 编辑部·[4 个来源]·
arXiv 上发表了四篇新的研究论文,探讨了老虎机算法的进展,重点关注反馈延迟、容量限制和内存集成等挑战。这些研究引入了新颖的算法和分析,以改进各种设置下的遗憾界限,包括具有任意延迟的 Lipschitz 老虎机、容量受限优化以及具有内存的广义线性老虎机。其中一篇论文专门讨论了线性上下文老虎机中的顺序批量学习,为个性化决策问题提供了更细粒度的表述。
AI
arXiv:2608.15036v1 Announce Type: new Abstract: The Lipschitz bandit problem extends the traditional multi-armed bandit framework to continuous action spaces by assuming that the reward functions satisfy a Lipschitz condition. This work investigates Lipschitz bandits under arbitr…
arXiv cs.LG
TIER_1English(EN)·Anling Xiang, Yuwen Yang, Yang Shen·
arXiv:2608.16216v1 Announce Type: new Abstract: What is the right delay complexity when a learner can track only $C$ pending feedback items and discarded feedback is permanently lost? Existing one-point bandit convex optimization guarantees in this model pay $\sqrt{T\sigma_{\max}…
arXiv:2608.15848v1 Announce Type: new Abstract: We study generalized linear bandits with memory, an endogenous non-stationary setting in which rewards depend on past actions through a finite memory matrix. Building on prior work for linear models (Clerici et al., 2024), we show t…
arXiv stat.ML
TIER_1English(EN)·Yanjun Han, Zhengqing Zhou, Zihao Hu, Jose Blanchet, Peter W. Glynn, Yinyu Ye, Zhengyuan Zhou·
arXiv:2004.06321v2 Announce Type: replace-cross Abstract: We study the sequential batch learning problem in linear contextual bandits with finite action sets, where the decision maker is constrained to split incoming individuals into (at most) a fixed number of batches and can on…