PulseAugur
实时 08:24:51
English(EN) The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models

新研究质疑语言模型中的叠加现象

一篇新论文分析了语言模型中“叠加”现象,即多个解决方案可能同时存在于单个表示中。研究人员使用三种不同的训练模式进行了调查:无训练、微调和从头开始训练。他们的发现表明,只有完全从头开始训练的模型才表现出利用叠加的迹象。相比之下,在无训练和微调模式下的模型要么崩溃了叠加,要么没有使用它,而是选择了捷径解决方案。 AI

影响 这项研究阐明了语言模型在复杂推理任务中可能利用或不利用叠加的条件。

排序理由 该集群包含一篇分析语言模型中特定现象的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新研究质疑语言模型中的叠加现象

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Michael Rizvi-Martel, Guillaume Rabusseau, Marius Mosbach ·

    The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models

    arXiv:2604.06374v2 Announce Type: replace Abstract: Latent reasoning via continuous chain-of-thoughts (Latent CoT) has emerged as a promising alternative to discrete CoT reasoning. Operating in continuous space increases expressivity and has been hypothesized to enable superposit…