PulseAugur
中
实时 04:25:59
English(EN) Optimal and Efficient Contextual Combinatorial Semi-bandits with General Function Approximation

新算法以更高效率解决上下文组合半老虎机问题

研究人员为上下文组合半老虎机问题开发了新算法,该问题涉及选择臂的子集以最大化累积奖励。其中一种方法,在最近的 arXiv 论文中提出,通过解决一个凸优化问题,提供了一种计算高效的方法来平衡探索和利用。该算法实现了 minimax 最优遗憾界限,并推广到任意组合动作结构和奖励函数逼近。另一篇论文侧重于预言机高效框架,显著减少了这些问题所需的预言机查询次数,特别是在最坏情况下的线性奖励设置中,同时保持了严格的遗憾保证。 AI

影响 这些在老虎机算法方面的进展可能导致在需要具有部分反馈的顺序选择的系统中,如推荐引擎或资源分配,实现更高效的决策。

排序理由 该集群包含两篇在 arXiv 上发表的学术论文,详细介绍了组合半老虎机问题的新算法和理论保证。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

新算法以更高效率解决上下文组合半老虎机问题

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含两篇在 arXiv 上发表的学术论文,详细介绍了组合半老虎机问题的新算法和理论保证。
Source corroboration
4 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
78 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [4]

  1. arXiv cs.LG TIER_1 English(EN) · Hao Qin, Chicheng Zhang ·

    具有通用函数逼近的 Optimal and Efficient Contextual Combinatorial Semi-bandits

    arXiv:2607.13686v1 Announce Type: new Abstract: We study the contextual combinatorial semi-bandit (CCSB) problem with general reward function approximation. At each round, the learner observes a context, selects a combinatorial action consisting of a subset of basic arms, and rec…

  2. arXiv cs.LG TIER_1 English(EN) · Chicheng Zhang ·

    具有通用函数逼近的最优高效上下文组合半赌博机

    We study the contextual combinatorial semi-bandit (CCSB) problem with general reward function approximation. At each round, the learner observes a context, selects a combinatorial action consisting of a subset of basic arms, and receives the reward of each selected arm; the goal …

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    具有通用函数逼近的最优高效上下文组合半老虎机

    We study the contextual combinatorial semi-bandit (CCSB) problem with general reward function approximation. At each round, the learner observes a context, selects a combinatorial action consisting of a subset of basic arms, and receives the reward of each selected arm; the goal …

  4. arXiv stat.ML TIER_1 English(EN) · Jung-hun Kim, Milan Vojnovi\'c, Min-hwan Oh ·

    Oracle-高效组合半老虎机

    arXiv:2510.21431v2 Announce Type: replace Abstract: We study the combinatorial semi-bandit problem where an agent selects a subset of base arms and receives individual feedback. While this generalizes the classical multi-armed bandit and has broad applicability, its scalability i…