PulseAugur
中
实时 04:14:56

新的CorrGRPO方法增强了语言模型的多元奖励学习

研究人员推出了一种名为相关性归一化GRPO(CorrGRPO)的新方法,用于训练具有多个奖励信号的推理语言模型。这种新方法解决了标准组相对策略优化(GRPO)中大规模奖励可能掩盖较小奖励的局限性。CorrGRPO将成对协方差归一化为皮尔逊相关系数,确保不同奖励成分的影响更加均衡。该方法已在0.5B到8B参数的模型上,在代码生成、工具调用和代理安全任务中展现出改进。 AI

影响 提高了多奖励语言模型的训练效率和性能,可能带来更强大的AI代理。

排序理由 该集群描述了研究论文中提出的一种用于训练语言模型的新方法。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的CorrGRPO方法增强了语言模型的多元奖励学习

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了研究论文中提出的一种用于训练语言模型的新方法。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
9 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Wenbin Hu, Huihao Jing, Haochen Shi, Yuxuan Liu, Haoran Li, Yangqiu Song ·

    CorrGRPO:用于多奖励学习的相关性归一化GRPO

    arXiv:2609.36820v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    CorrGRPO:用于多奖励学习的相关性归一化GRPO

    Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward…