English(EN)Decoding Looped Transformers Better for (Almost) Free
新研究探索Transformer模型的效率和稳定性
作者PulseAugur 编辑部·[9 个来源]·
研究人员正在探索新的方法来提高Transformer模型的效率和性能,特别是那些采用循环的Transformer模型。其中一种方法LoopCD提供了一个无需训练的框架,通过对比最终预测与早期循环传递来提高解码质量,在代码生成和推理等任务中取得了显著的进步。另一个研究领域侧重于理解和稳定深度图Transformer,分析其动力学系统以防止表示崩溃并改进图生成。此外,研究正在调查循环Transformer如何在共享权重中路由计算,表明中间状态和学习到的引导层控制着执行的具体操作。最后,研究正在检查条件功能可替代性概念,以理解Transformer中的冗余和缩放,揭示性能提升并不总是与可替代性增加相关,并提出了高效模型缩放的新方向。
AI
arXiv:2610.02185v1 Announce Type: new Abstract: Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlie…
arXiv:2609.39739v1 Announce Type: new Abstract: Graph generative models increasingly rely on Graph Transformers (GT) to capture complex dependencies among nodes and edges. While deeper architectures should provide greater expressive capacity and a broader receptive field, their e…
arXiv cs.LG
TIER_1English(EN)·Jiaju Wu, Yi Hu, Muhan Zhang·
arXiv:2609.39892v1 Announce Type: new Abstract: Looped Transformers repeatedly apply the same set of Transformer layers, giving them a recurrent architecture for latent computation. Their strong performance on iterative reasoning and length-generalization tasks suggests an appeal…
arXiv cs.AI
TIER_1English(EN)·Jiaheng Chen, Jiaxing Li, Yucheng Xiao, Xinyong Cai, Juncheng Bu, Lan Yu, Tinghe Zhang·
arXiv:2609.39259v1 Announce Type: new Abstract: Modern neural networks scale predictably, yet the mechanisms behind these regularities remain unclear. Neural redundancy is typically characterized by component importance or representational similarity, both indirect proxies. We vi…
Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less comp…
arXiv:2609.36653v1 Announce Type: new Abstract: Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with shared parameters. However, these models apply each learned update with a fixed unit…
arXiv:2609.37379v1 Announce Type: new Abstract: Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models have likewise highlighted the value of scaling test-time computation through longe…
Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with shared parameters. However, these models apply each learned update with a fixed unit scale, which can be conservative when updates m…
<p><b><span style="white-space: pre-wrap;">TL;DR</span></b><span style="white-space: pre-wrap;">: </span></p><ul><li value="1"><span style="white-space: pre-wrap;">We identify a confound in no-cot-bench related to the positioning of the prompt's "key." </span></li><li value="2"><…