PulseAugur
中
实时 01:27:16
English(EN) The Residual Stream's Effective Depth

新指标揭示 Transformer 模型如何利用深度

研究人员引入了一个名为“有效深度”($\Deff$)的新指标来分析 Transformer 语言模型的残差流。该指标量化了表示相似性随层距离衰减的程度,并提供一个单一的标量值。在十六个仅解码器模型中,$\Deff$ 显示大多数模型表现出相关残差更新的校准特征,而不是表明深度未被使用。值得注意的是,Qwen3.5 和 OLMo-2 的有效深度显著低于其理论最大值,这表明其残差流处理存在特定模式。 AI

影响 引入了一种理解模型内部动态的新诊断工具,可能有助于未来的模型开发和分析。

排序理由 学术论文,介绍用于分析 Transformer 模型的新指标。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新指标揭示 Transformer 模型如何利用深度

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Barak Gahtan, Ido Galil, Alex M. Bronstein ·

    残差流的有效深度

    arXiv:2609.31098v1 Announce Type: new Abstract: We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggreg…