PulseAugur
中
实时 13:27:30
English(EN) The Authenticity Gap in Human Evaluation

新论文质疑了NLG模型的标准人类评估方法

一篇新发表在arXiv上的论文批评了自然语言生成(NLG)系统的标准人类评估协议。作者认为,常见的做法,特别是使用李克特量表,可能导致对人类偏好的评估不准确,甚至会颠倒真实的偏好方向。他们提出了一种名为系统级概率评估(SPA)的替代方法,用于评估故事生成等开放式任务,并证明了其在正确排序GPT-3模型大小方面的有效性。 AI

影响 提出了一种新的评估协议,可能导致对NLG模型能力进行更准确的评估。

排序理由 发表在arXiv上的研究论文,详细介绍了NLG模型的新评估协议。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新论文质疑了NLG模型的标准人类评估方法

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
发表在arXiv上的研究论文,详细介绍了NLG模型的新评估协议。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
50 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Kawin Ethayarajh, Dan Jurafsky ·

    人类评估中的真实性差距

    arXiv:2205.11930v3 Announce Type: replace-cross Abstract: Human ratings are the gold standard in NLG evaluation. The standard protocol is to collect ratings of generated text, average across annotators, and rank NLG systems by their average scores. However, little consideration h…