PulseAugur
实时 11:01:11
English(EN) Whose Gold? Annotator-Pool Disagreement Is Large at the Item Level, and Hidden by Small Leaderboards

专家与众包标注者之间的分歧掩盖了模型性能的变异性

arXiv上最近发表的一项研究强调了专家和众包标注者在评估语言模型时存在显著分歧。尽管在项目层面的判断上存在这些实质性差异,但由此产生的模型排行榜却基本相同,这表明当前的聚合方法掩盖了潜在的变异性。研究表明,标注者池的选择会影响模型的可感知性能,并且LLM裁判比专家更倾向于与众包标注者保持一致。 AI

影响 强调了人工智能模型评估中潜在的偏见以及对更稳健的基准测试方法的需求。

排序理由 该集群包含一篇发表在arXiv上的研究论文,讨论了人工智能模型评估的方法论。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

专家与众包标注者之间的分歧掩盖了模型性能的变异性

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Anik Jha ·

    谁的金子?标注者池在项目层面的分歧很大,被小型排行榜掩盖

    arXiv:2608.15980v1 Announce Type: new Abstract: Preference benchmarks are built by hiring annotators, and the identity of those annotators is treated as an implementation detail. We measure what that detail buys. On the 2,885 MultiPref items where both pools are internally unanim…