PulseAugur
中
实时 20:29:48
English(EN) acc vs acc_norm: Why Length Bias Skews LLM Eval Scores

LLM评估指标'acc'与'acc_norm'详解

一项技术性讨论强调了在评估大型语言模型(LLM)时,“acc vs acc_norm”问题,尤其是在多项选择任务中。标准准确率指标(acc)由于依赖于总对数似然,而总对数似然本身会因答案长度而受到惩罚,因此偏向于较短的答案。归一化准确率指标(acc_norm)试图通过将分数除以续写的字节长度来纠正这一点,旨在实现跨模型和分词器的一致性比较。然而,acc_norm也可能引入偏差,例如偏向于将答案拆分成许多可预测的子词的答案。文章建议报告这两个指标,并在训练前为每个任务系列选择一个一致的指标,以避免误导性评估。 AI

影响 阐明了LLM评估指标中潜在的偏差,指导研究人员进行更准确的模型比较。

排序理由 讨论LLM评估指标的技术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

LLM评估指标'acc'与'acc_norm'详解

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
讨论LLM评估指标的技术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
46 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. dev.to — LLM tag TIER_1 English(EN) · jidonglab ·

    acc vs acc_norm:长度偏差如何影响 LLM 评估分数

    <p>Your fine-tune gains three points of <code>acc_norm</code> on HellaSwag and loses two points of <code>acc</code>. Same checkpoint, same harness, same seed. Nothing about the model's commonsense reasoning moved in two directions at once — you changed its average per-token entro…