PulseAugur
EN
LIVE 00:04:46

New metric reveals how transformer models use depth

Researchers have introduced a new metric called 'effective depth' ($\Deff$) to analyze the residual stream in transformer language models. This metric quantifies how representation similarity decays with layer distance, providing a single scalar value. Across sixteen decoder-only models, $\Deff$ revealed that most models exhibit a calibrated signature of correlated residual updates rather than indicating unused depth. Notably, Qwen3.5 and OLMo-2 showed significantly lower effective depths compared to their theoretical maximums, suggesting a specific pattern in their residual stream processing. AI

IMPACT Introduces a novel diagnostic tool for understanding internal model dynamics, potentially aiding in future model development and analysis.

RANK_REASON Academic paper introducing a new metric for analyzing transformer models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.LG →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New metric reveals how transformer models use depth

COVERAGE [1]

  1. arXiv cs.LG TIER_1 English(EN) · Barak Gahtan, Ido Galil, Alex M. Bronstein ·

    The Residual Stream's Effective Depth

    arXiv:2609.31098v1 Announce Type: new Abstract: We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggreg…