PulseAugur
中
实时 15:03:43
English(EN) Low-Rank Friction for Memory-Efficient Transformer Pretraining

新的 Rank-1 iKFAD 优化器将 Transformer 预训练的内存减半

研究人员推出了一种新颖的优化技术 Rank-1 iKFAD (R-iKFAD),旨在提高 Transformer 预训练的内存效率。该方法通过将 iKFAD 优化器的完整摩擦张量替换为秩为 1 的外积分解,将每层的内存占用从 O(mn) 显著降低到 O(m+n)。在 GPT2-Nano 和 DistilBERT 等各种模型上的实验表明,R-iKFAD 在性能上与原始 iKFAD 相当,同时将优化器的内存需求几乎减半。 AI

影响 这种内存高效的优化技术可以使在现有硬件上训练更大的 Transformer 模型成为可能,从而可能加速研究和开发。

排序理由 该集群包含一篇详细介绍 Transformer 模型新优化技术的论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv stat.ML 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的 Rank-1 iKFAD 优化器将 Transformer 预训练的内存减半

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍 Transformer 模型新优化技术的论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
2 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv stat.ML TIER_1 English(EN) · Rajit Rajpal, Benedict Leimkuhler ·

    用于内存高效 Transformer 预训练的低秩摩擦

    arXiv:2609.30342v1 Announce Type: new Abstract: iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor $\xi\in\mathbb{R}^{m\times n…