PulseAugur
中
实时 12:40:14
English(EN) CERO: Where and When to Allocate Rollouts for RL Post-Training

新的CERO调度器优化RL后训练推出预算

研究人员推出了一种新颖的在线对偶规划调度器CERO,旨在优化强化学习(RL)后训练的推出预算分配。与固定每次更新预算的传统方法不同,CERO通过调整提示选择、重访频率和组生成,在整个训练过程中协调有限的预算。该方法利用紧凑的Fenchel表示和投影在线梯度下降,在数学推理基准上的表现优于固定速率和时变基准。 AI

影响 优化强化学习的资源分配,可能提高复杂推理任务的训练效率和性能。

排序理由 这是一篇详细介绍强化学习新方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的CERO调度器优化RL后训练推出预算

本文如何被排名

Signal score
8 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是一篇详细介绍强化学习新方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Yiming Zong, Yige Wang, Xing Hu, Jiashuo Jiang, Zuo-Jun Max Shen ·

    CERO:在哪里以及何时为 RL 后训练分配部署

    arXiv:2610.09679v1 Announce Type: new Abstract: Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to coordinate a finite rollout budget over the entire training horizon. We formulat…