PulseAugur
实时 21:48:17
English(EN) 356 of the 528 eval cases we could judge have never failed. Our pass rate could only move 21 points.

LLM 评估套件因大多数测试用例停止区分而失效

LLM 评估套件的分析显示,其大部分测试用例随着时间的推移已变得无效。在具有足够版本历史的 528 个案例中,有 356 个在所有测试版本中始终通过,59 个始终失败,表明它们没有提供区分能力。只有 113 个案例(占判断集的 21%)在不同版本之间确实存在差异,这表明随着错误的修复且未重新引入,该套件检测回归的能力已经减弱。 AI

影响 凸显了为快速发展的 LLM 维护有效评估套件的挑战。

排序理由 对 LLM 评估套件历史性能的分析。[lever_c_demoted from research: ic=1 ai=1.0]

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

LLM 评估套件因大多数测试用例停止区分而失效

报道来源 [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Ethan Walker ·

    356 of the 528 eval cases we could judge have never failed. Our pass rate could only move 21 points.

    <p>TL;DR: I broke out per-case results for the incident-harvested part of our eval suite, 611 of our 1,400 cases, and asked which cases have ever discriminated between two shipped versions. Of the 528 with enough history to judge, 356 had passed every version and 59 had failed ev…