PulseAugur
实时 10:31:42
English(EN) Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

研究发现LLM评估排名随令牌预算变化

arXiv上的一篇新研究论文探讨了大型语言模型的评估如何受到令牌生成预算的显著影响。研究发现,在3-19%的情况下,随着预算的增加,模型的准确性会下降,并且这些影响因模型而异。在所有测试的基准上,模型排名在不同预算下都会发生逆转,这表明标准评估可能无法准确反映真实性能。研究还强调了模型之间潜在的互补性,并提出一个预算感知的路由系统可以捕捉到部分性能差距。 AI

影响 强调需要更细致的LLM评估协议,以考虑不同的计算预算。

排序理由 关于LLM评估方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究发现LLM评估排名随令牌预算变化

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Rodrigo Guedes de Souza, Alison R. Panisson ·

    谁思考得更好取决于你给他们多长时间:LLM评估中的预算依赖排名

    arXiv:2608.12150v1 Announce Type: new Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven …