PulseAugur
中
实时 08:55:28
English(EN) Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification Signals

新研究探讨大型语言模型检测无法回答问题的能力

研究人员开发了一种方法来检测大型语言模型(LLM)何时在回答它们无法真正回答的问题,或在对话中过早回应。他们创建了一个新的基准和评估工具,以在六个数据集和六个开源LLM上测试此能力。研究结果表明,无法回答问题的信号在相似数据集之间转移良好,例如涉及数学问题或文本段落中信息缺失的情况,但在转移到其他类型的无法回答问题(如认识论上的“已知未知”)时效果不佳。虽然经过校准的探测器可以在不进行模型微调的情况下准确识别欠指定的回合,但其在最终任务上的成功有限,这表明剩余的差距在于模型如何利用澄清而不是检测。 AI

影响 这项研究可能有助于开发出更可靠地识别和拒绝回答无法回答问题的LLM,从而改进对话系统和信息检索。

排序理由 该集群包含一篇详细介绍LLM能力评估新方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新研究探讨大型语言模型检测无法回答问题的能力

本文如何被排名

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍LLM能力评估新方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Danil Fedorov, Kirill Redko, Sergey Chuprin, Aidar Shumbalov, Stanislav Chumakov, Anna Kalyuzhnaya ·

    知道何时不回答:潜在欠定信号的跨域和多轮泛化

    arXiv:2610.08413v1 Announce Type: cross Abstract: Large language models routinely answer questions that cannot be answered from the information given, and in dialogue they answer before enough has been said. Unanswerability is linearly decodable from hidden states, but it is uncl…