PulseAugur
实时 05:49:19
English(EN) What Is the Minimum Architecture for Prolepsis? Early Irrevocable Commitment Across Tasks in Small Transformers

新“prolepsis”现象在小型 Transformer 模型中被识别

研究人员在小型 Transformer 模型中识别出一种称为“prolepsis”的现象,即模型在处理早期就做出决定,且无法纠正。这种承诺由特定任务的注意力头维持,并且不易被标准的残差流方法检测到,尽管基于 CLT 的引导显示出一些成功。研究发现,这种 prolepsis 模式出现在 Gemma 2-2BLlama 3.2 1B 等仅解码器模型中的不同任务上,表明存在一个共享的潜在机制。 AI

影响 识别出小型 Transformer 模型的一个新局限性,可能影响其可靠性和可解释性。

排序理由 该集群包含一篇详细介绍 Transformer 模型中观察到的新现象的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新“prolepsis”现象在小型 Transformer 模型中被识别

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · \'Eric Jacopin ·

    什么是预示的最小架构?小型 Transformer 模型在任务中的早期不可撤销承诺

    arXiv:2604.15010v2 Announce Type: replace-cross Abstract: When do transformers commit to a decision, and what prevents them from correcting it? We introduce prolepsis: a transformer commits early, task-specific attention heads sustain the commitment, and no layer corrects it. Rep…