PulseAugur
实时 10:58:00
English(EN) Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap

研究发现:Token上限扭曲了多语言人工智能推理测试

麦格理商学院的一篇新研究论文调查了多语言评估中的输出Token上限如何扭曲结果。研究发现,多语言推理中测得的差距,特别是对于德语、泰语和斯瓦希里语等语言,根据使用的Token预算的不同,可能存在显著差异(高达57个点)。研究人员证明,调整输出上限或使用长度归一化可以改变性能指标,这表明当前的评估方法可能无法准确反映模型在不同语言中的真实能力。 AI

影响 强调了当前多语言LLM评估中潜在的偏见,表明需要修订测试方法。

排序理由 学术论文,详细介绍了评估LLM性能的新颖方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究发现:Token上限扭曲了多语言人工智能推理测试

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Ankit Goyal, Jaideep Ray ·

    注意上限:输出预算机制改变了测量的多语言推理差距

    arXiv:2608.04160v1 Announce Type: new Abstract: Multilingual evaluations report accuracy at a single output-token cap, but languages need different numbers of tokens to express the same content, so the cap is a hidden experimental variable. We test whether the native-vs-translate…