PulseAugur
实时 00:43:39
English(EN) Every eval you run against a public benchmark is a training signal you hand the next model. The leaderboard isn't measuring capability, it's leaking answers int

AI基准评估存在污染未来模型训练数据的风险

在公开基准上运行评估会无意中训练未来的AI模型,因为这些评估会成为预训练数据的一部分。这种污染意味着随着时间的推移,排行榜可能反映的是学到的答案,而不是真实的能力。 AI

影响 引发了对当前AI基准有效性和未来模型开发完整性的质疑。

排序理由 该条目讨论的是AI评估方法论的一个概念性问题,而不是一个具体的事件或发布。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

AI基准评估存在污染未来模型训练数据的风险

报道来源 [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    你对公开基准进行的每一次评估都是你交给下一个模型的训练信号。排行榜没有衡量能力,它正在泄露答案

    Every eval you run against a public benchmark is a training signal you hand the next model. The leaderboard isn't measuring capability, it's leaking answers into the pretraining set. Contamination isn't a flaw in the score — after enough cycles, it IS the score. # AI # MachineLea…