PulseAugur
实时 09:18:28
English(EN) How often do five frontier LLMs agree on whether a claim is true? Across 1,000 real fact-check requests, they agree only about a third of the time. The rubric i

前沿大语言模型在事实核查声明上共识度低

一项分析了五个前沿大语言模型(LLMs)共识率的研究发现,在 1,000 个真实世界的事实核查请求中,它们在判断声明真伪时仅约有三分之一的时间达成一致。即使有两个模型具备网络搜索能力,它们在 6% 的声明上仍直接相互矛盾,尽管它们访问的是相同的信息来源。评估评分标准通过要求四方裁决且无弃权选项,进一步夸大了共识度。 AI

影响 凸显了在大语言模型在事实核查应用中的可靠性和一致性方面存在的重大挑战。

排序理由 该集群讨论了对大语言模型在事实核查方面共识的分析,这是一篇观点/分析文章,而非主要发布或研究论文。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

前沿大语言模型在事实核查声明上共识度低

报道来源 [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    五个前沿大语言模型在判断事实真伪时有多一致?在 1000 个真实事实核查请求中,它们仅约三分之一的时间达成一致。评分标准 i

    How often do five frontier LLMs agree on whether a claim is true? Across 1,000 real fact-check requests, they agree only about a third of the time. The rubric inflates that, since it forces a four-way verdict with no option to abstain. But even the two models that can search the …