PulseAugur
实时 09:36:56
English(EN) Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express

LLM内部编码因果推理能力,但无法口头表达

一篇题为“因果语言障碍”(Causal Tongue-Tie)的新研究论文指出,大型语言模型(LLMs)在因果推理方面的内部编码与其实际输出之间存在差异。研究发现,虽然LLMs可以在其隐藏状态中准确表示因果方向(通过线性探测达到约0.97的准确率),但它们直接的是/否回答常常回归常识,准确率仅为约0.5。这种“因果语言障碍”表明,LLM的输出可能无法完全反映其内部理解,这使得评估其因果推理能力变得复杂。 AI

影响 这项研究表明,当前评估LLM因果推理能力的基准可能存在缺陷,需要新的方法来探测内部状态,而不仅仅是输出。

排序理由 该集群包含一篇在arXiv上发表的研究论文,详细介绍了LLM的能力发现。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

LLM内部编码因果推理能力,但无法口头表达

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Ziyi Ding, Xiao-Ping Zhang ·

    因果性结巴:大型语言模型可以编码因果方向,但其是/否输出未能表达

    arXiv:2605.25891v1 Announce Type: cross Abstract: We find a mismatch between what large language models encode about a causal question and what they answer. On anti-commonsense CLadder items, a fixed linear probe recovers the evidence-supported answer from the model's hidden stat…

  2. arXiv cs.AI TIER_1 English(EN) · Xiao-Ping Zhang ·

    因果舌缠:大型语言模型能编码因果方向,但其是/否输出未能表达

    We find a mismatch between what large language models encode about a causal question and what they answer. On anti-commonsense CLadder items, a fixed linear probe recovers the evidence-supported answer from the model's hidden state (accuracy approximately 0.97), while the spoken …