PulseAugur
实时 09:12:38
English(EN) Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs

新研究探讨大型语言模型在不同语言和任务中的不确定性估计 · 跟踪 4 个来源

研究人员正在探索提高大型语言模型 (LLM) 在各种语言和任务中的不确定性估计的方法。一项研究发现,即使问题是低资源语言,提示 LLM 用英语进行推理也能显著提高不确定性估计性能。另一篇论文提出了一个框架,将 LLM 的不确定性分解为输入歧义、知识差距和解码随机性,从而为审计可靠性提供更细致的理解。此外,一种新方法使用知识蒸馏来创建高效的、单通道的 LLM 进行不确定性估计,其性能与计算密集型方法相当。 AI

影响 这些研究旨在通过使 LLM 能够识别和量化自身的不确定性来提高其可靠性和可信度,这对于在关键应用中安全部署至关重要。

排序理由 该集群包含多篇在 arXiv 上发表的关于 LLM 不确定性估计的学术论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

新研究探讨大型语言模型在不同语言和任务中的不确定性估计 · 跟踪 4 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含多篇在 arXiv 上发表的关于 LLM 不确定性估计的学术论文。
Source corroboration
4 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
58 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [4]

  1. arXiv cs.AI TIER_1 English(EN) · Andrea Alfarano, Andrea Bacciu, Saab Mansour, Amin Mantrach, Marcello Federico ·

    从推理中估计不确定性:LLM多语种和跨语种MCQA性能的大规模研究

    arXiv:2607.06327v1 Announce Type: cross Abstract: Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across 22 languages, spann…

  2. arXiv cs.AI TIER_1 English(EN) · Marcello Federico ·

    从推理中估计不确定性:LLM多语种和跨语种MCQA性能的大规模研究

    Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across 22 languages, spanning high-, mid-, and low-resource settings. Using …

  3. arXiv cs.AI TIER_1 English(EN) · Aditya Taparia, Ransalu Senanayake, Kowshik Thopalli, Vivek Narayanaswamy ·

    LLM 中不确定性的解剖

    arXiv:2603.24967v2 Announce Type: replace Abstract: Understanding why a large language model (LLM) is uncertain about the response is important for their reliable deployment. Current approaches, which either provide a single uncertainty score or rely on the classical aleatoric-ep…

  4. arXiv stat.ML TIER_1 English(EN) · Lakshmana Sri Harsha Nemani, P. K. Srijith, Tomasz Ku\'smierczyk ·

    通过证据知识蒸馏实现 LLM 中高效的不确定性

    arXiv:2507.18366v2 Announce Type: replace-cross Abstract: Accurate uncertainty quantification remains a key challenge for standard LLMs, prompting the adoption of Bayesian and ensemble-based methods. However, such methods typically necessitate computationally expensive sampling, …