PulseAugur
实时 07:16:01
English(EN) ScalePRM: Training Process Reward Models by Scaling Verification Compute Without Ground Truth

ScalePRM 在无需真实标签的情况下训练 AI 奖励模型,性能优于 GPT-4o

研究人员开发了 ScalePRM,一种新颖的训练过程奖励模型(PRMs)的方法,该方法绕过了昂贵的步骤级正确性标签或真实标签的需要。该方法通过生成和聚合每个推理步骤的多个独立验证来创建合成标签,从而扩展了验证计算。ScalePRM 在 ProcessBench 基准测试上取得了 67.5 的 F1 分数,性能优于传统的基于真实标签的训练,甚至优于作为批评者的 GPT-4o。当用作 Qwen2.5-Math-7B 的 RL 训练的奖励信号时,它将六个数学推理基准的平均准确率提高到 47.4%,超过了基于真实标签的 RLVR。 AI

影响 这种方法可以显著降低需要逐步推理的 AI 模型训练的成本和复杂性,有可能加速数学问题解决等领域的进展。

排序理由 该集群描述了一篇详细介绍 AI 模型训练新方法的最新研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

ScalePRM 在无需真实标签的情况下训练 AI 奖励模型,性能优于 GPT-4o

本文如何被排名

Signal score
23 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一篇详细介绍 AI 模型训练新方法的最新研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Salman Rahman, Sruthi Gorantla, Arpit Gupta, Swastik Roy, Nanyun Peng, Yang Liu ·

    ScalePRM:通过扩展验证计算而不使用真实标签来训练过程奖励模型

    arXiv:2512.03244v2 Announce Type: replace-cross Abstract: Training process reward models (PRMs) requires step-level correctness labels, obtained either through expensive human annotation or by relying on ground-truth answers, limiting the ability to scale process-level supervisio…