PulseAugur
中
实时 20:47:34
English(EN) Sequential Batch Learning in Finite-Action Linear Contextual Bandits

新研究探讨老虎机算法以改进决策和减少遗憾 · 跟踪 8 个来源

几篇近期研究论文探讨了老虎机算法的进展,这是一种顺序决策框架。一篇论文介绍了潜在顺序老虎机(LOB),它通过仅要求了解状态内动作偏好的部分顺序来放宽先前潜在老虎机算法的假设,从而提高样本效率。另一项研究侧重于多臂老虎机中遗憾与不稳定性之间的权衡,提出了一种新的算法 SLE-UCB,该算法匹配理论下界。进一步的研究解决了具有任意反馈延迟的 Lipschitz 老虎机,开发了实现强遗憾保证的算法,并探讨了容量受限延迟老虎机优化中的条件能量和时间几何,揭示了基于时间和容量的遗憾的细微差别。最后,一篇关于上下文老虎机的论文提出了一种快速、同类最佳的遗憾算法,另一篇则研究了线性上下文老虎机中的顺序批量学习,为实际应用提供了近乎完整的表征。 AI

影响 老虎机算法的这些进展可能导致在从个性化到复杂优化问题的各种应用中,AI 系统的决策更加高效和有效。

排序理由 多篇 arXiv 论文详细介绍了老虎机优化领域的新理论算法和分析。

在 arXiv stat.ML 阅读 →

AI 生成摘要 · Google Gemini · 来自 8 个来源。 我们如何撰写摘要 →

新研究探讨老虎机算法以改进决策和减少遗憾 · 跟踪 8 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇 arXiv 论文详细介绍了老虎机优化领域的新理论算法和分析。
Source corroboration
8 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
46 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [8]

  1. arXiv cs.LG TIER_1 Deutsch(DE) · Emil Carlsson, Newton Mwai, Fredrik D. Johansson ·

    Latent Order Bandits

    arXiv:2605.07304v2 Announce Type: replace Abstract: Bandit algorithms solve diverse sequential decision-making problems, but are often too sample-inefficient for from-scratch personalization. To substantially reduce exploration times, latent bandit algorithms exploit cross-instan…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    多臂老虎机中走向最优遗憾-不稳定性权衡

    Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs. We study the trade-off between worst-case regret $\mathcal{R}_{K,T}$ and instability $\mathcal S_{K,T}$, defined as the largest standard de…

  3. arXiv cs.LG TIER_1 English(EN) · Yuhao Liu, Yu Chen, Longbo Huang ·

    具有任意延迟的Lipschitz Bandit

    arXiv:2608.15036v1 Announce Type: new Abstract: The Lipschitz bandit problem extends the traditional multi-armed bandit framework to continuous action spaces by assuming that the reward functions satisfy a Lipschitz condition. This work investigates Lipschitz bandits under arbitr…

  4. arXiv cs.LG TIER_1 English(EN) · Anling Xiang, Yuwen Yang, Yang Shen ·

    超越峰值积压:容量受限延迟老虎机优化中的条件能量与时间几何

    arXiv:2608.16216v1 Announce Type: new Abstract: What is the right delay complexity when a learner can track only $C$ pending feedback items and discarded feedback is permanently lost? Existing one-point bandit convex optimization guarantees in this model pay $\sqrt{T\sigma_{\max}…

  5. arXiv stat.ML TIER_1 English(EN) · Samuel Girard, Aurelien Bibaut, Arthur Gretton, Nathan Kallus, Houssam Zenati ·

    Contextual Bandits 的快速最佳遗憾

    arXiv:2510.15483v3 Announce Type: replace Abstract: We study the problem of stochastic contextual bandits in the agnostic setting, where the goal is to compete with the best policy in a given class without assuming realizability or imposing model restrictions on losses or rewards…

  6. arXiv stat.ML TIER_1 English(EN) · Kaifei Wang, Yinyu Ye, Han Zhong ·

    迈向多臂老虎机中的最优遗憾-不稳定性权衡

    arXiv:2608.17841v1 Announce Type: new Abstract: Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs. We study the trade-off between worst-case regret $\mathcal{R}_{K,T}$ and instability $\math…

  7. arXiv stat.ML TIER_1 English(EN) · Heesang Ann, Hyunjun Choi, Taehyun Hwang, Younghoon Shin, Haeju Cheong, Min-hwan Oh ·

    具有记忆的广义线性老虎机

    arXiv:2608.15848v1 Announce Type: new Abstract: We study generalized linear bandits with memory, an endogenous non-stationary setting in which rewards depend on past actions through a finite memory matrix. Building on prior work for linear models (Clerici et al., 2024), we show t…

  8. arXiv stat.ML TIER_1 English(EN) · Yanjun Han, Zhengqing Zhou, Zihao Hu, Jose Blanchet, Peter W. Glynn, Yinyu Ye, Zhengyuan Zhou ·

    有限动作线性上下文老虎机中的顺序批量学习

    arXiv:2004.06321v2 Announce Type: replace-cross Abstract: We study the sequential batch learning problem in linear contextual bandits with finite action sets, where the decision maker is constrained to split incoming individuals into (at most) a fixed number of batches and can on…