PulseAugur
中
实时 16:52:56
English(EN) Learning Rate Transfer for Hybrid Transformer-SSM Architectures

新研究揭示混合Transformer-SSM模型中的学习率迁移

一篇新研究论文探讨了混合架构的学习率(LR)缩放问题,该架构结合了Transformer和状态空间模型(SSM)模块,这些模块越来越多地用于生产语言模型中。研究发现,这些混合模型的实际实现,即使使用了简化的SSM,在各种宽度和深度下(高达十亿参数规模)也表现出接近零的学习率迁移差距。这种不变性归因于$\mu$P的初始化和学习率缩放,以及AdamW的逐参数归一化,它们共同维持了全局更新到权重的 Invariance 和局部组件的平衡。 AI

影响 这项研究通过阐明混合架构中的学习率迁移,有望提高大型语言模型的训练效率和可扩展性。

排序理由 该集群包含一篇详细介绍模型架构和训练方法新研究发现的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新研究揭示混合Transformer-SSM模型中的学习率迁移

本文如何被排名

Signal score
4 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍模型架构和训练方法新研究发现的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Jimin Seo, Gyubok Lee, Yeonsik Jo, Kiwoong Yoo, Yeongoon Kim, Minhae Oh, Jin Woo Koo, Suhwan Kim, Nakyung Lee, Minsik Seol, Idris Nechnech, Jaehyeon Kim, Giho Lee, Jungwoo Lee ·

    混合 Transformer-SSM 架构的学习率迁移

    arXiv:2610.01172v1 Announce Type: new Abstract: We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models. In particular, we focus on the gap between the theo…