PulseAugur
实时 09:44:35

研究发现:LLM在多轮证据收集推理方面存在困难

一篇题为《Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference》的新论文介绍了一个“外星人绑架游戏”,用于测试大型语言模型(LLMs)在多轮溯因推理场景中获取证据和完善假设的能力。研究发现,与在多轮中提供证据相比,一次性提供所有证据时LLMs的表现更好。此外,接收预言家提供示例的模型比自行选择查询的模型成功率更高,尽管它们最终的假设与自行选择的证据更一致。 AI

影响 这项研究突显了LLM随着时间推移主动寻求和整合信息的能力的局限性,暗示了在复杂推理任务中可能存在的问题。

排序理由 该集群包含一篇详细介绍LLM能力研究结果的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究发现:LLM在多轮证据收集推理方面存在困难

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Shahrukh Mohiuddin, Chalamalasetti Kranti, Sherzod Hakimov, David Schlangen ·

    Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference

    arXiv:2608.03388v1 Announce Type: new Abstract: Abductive reasoning requires forming hypotheses that explain observed evidence and revising them as new evidence becomes available. While large language models (LLMs) are often evaluated on whether they solve abductive reasoning tas…