PulseAugur
实时 13:22:37
English(EN) Germany's Soofi consortium withdrew benchmark scores after discovering its training data contained paraphrased questions from the GPQA evaluation set. The 11.1-

Soofi联盟因数据污染撤回GPQA基准分数

德国Soofi联盟已撤回其基准分数,原因是发现其训练数据包含GPQA评估集的释义问题。这种污染意味着之前报告的GPQA-Diamond基准11.1分的提升不再是衡量模型能力的有效指标。一名研究人员发现了这个问题,Soofi团队在一周内确认了数据污染。 AI

影响 训练数据中的污染会使基准测试结果无效,这凸显了在AI研究中进行严格评估和数据卫生的必要性。

排序理由 因研究评估中的数据污染而撤回基准分数。[lever_c_research降级:ic=1 ai=1.0]

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Soofi联盟因数据污染撤回GPQA基准分数

报道来源 [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Germany's Soofi consortium withdrew benchmark scores after discovering its training data contained paraphrased questions from the GPQA evaluation set. The 11.1-

    Germany's Soofi consortium withdrew benchmark scores after discovering its training data contained paraphrased questions from the GPQA evaluation set. The 11.1-point gain on GPQA-Diamond can no longer stand as evidence of model capability. A researcher spotted the contamination; …