PulseAugur
中
实时 12:40:27
English(EN) Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL

新工具 BehaviorTrace 质疑强化学习模型中的训练数据归因

研究人员开发了 BehaviorTrace,这是一个用于评估在线强化学习(RL)中语言模型训练数据归因的开放评估工具。该研究在 Qwen2.5-1.5B 模型上使用 GRPO 进行,发现许多明显的归因信号实际上是混淆因素。一种简单的梯度幅度排序方法与目标估计器相比表现相当,模型流畅性也被证明是学习行为的有力预测指标。尽管每次部署的结果因种子和生成抽样而异,但出现了一个一致的信号:触发 token 的梯度与行为的实际发生一致。 AI

影响 强调了在强化学习模型中将学习行为归因于特定训练数据的局限性,表明需要更鲁棒的评估方法。

排序理由 研究论文,详细介绍了强化学习中归因方法的新评估工具。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新工具 BehaviorTrace 质疑强化学习模型中的训练数据归因

本文如何被排名

Signal score
8 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
研究论文,详细介绍了强化学习中归因方法的新评估工具。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Amit Nautiyal ·

    哪个推广教会了它?BehaviorTrace 与在线强化学习中训练数据归因的局限性

    arXiv:2610.10422v1 Announce Type: cross Abstract: When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it? And when an attribution method says it can, how do we know the answer is real? We study both questions on onlin…