PulseAugur
中
实时 08:49:18
English(EN) Memory-Efficient Expert Routing for Distributed MoE Training

RelayMoE 提升 MoE 训练效率和内存使用率

研究人员开发了 RelayMoE,这是一种新颖的基于环的执行模型,旨在提高混合专家(MoE)模型分布式训练期间的内存效率。该方法通过循环专家权重或令牌并在本地计算,避免了构建完整的 top-k 扩展调度缓冲区。RelayMoE 还可以实现内存高效的反向传播重计算,从而支持更长的序列和更大的批次,或者保留更多的注意力激活以提高训练吞吐量。在 30B-57B 参数 MoE 模型上的评估表明,与 Megatron-LM 相比,速度提高了 2 倍,在相同的内存限制下,吞吐量提高了 2.02 倍。 AI

影响 提出了一种显著提高大型 MoE 模型训练效率和内存使用率的方法,有望支持更大模型和更长上下文。

排序理由 学术论文,详细介绍了分布式 MoE 模型训练的新方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

RelayMoE 提升 MoE 训练效率和内存使用率

本文如何被排名

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
学术论文,详细介绍了分布式 MoE 模型训练的新方法。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Arnab Kanti Tarafder, Jaume Guasch-Mart\'i, Gokcen Kestor, Jie Ren ·

    分布式MoE训练的内存高效专家路由

    arXiv:2610.07333v1 Announce Type: cross Abstract: As Mixture-of-Experts (MoE) models scale toward hundreds of experts and higher top-$k$ routing, memory efficiency in distributed training becomes a critical bottleneck. Peak memory is dominated by the MoE block, not attention: eve…