PulseAugur
实时 01:56:13
English(EN) Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

新框架高效预测大型MoE模型最优学习率

研究人员开发了一种新颖的两步框架,可高效确定大型混合专家(MoE)模型的最优学习率。该方法利用不同模型宽度之间的超参数迁移,并将研究结果外推到海量token预算,显著降低了传统超参数搜索的计算成本。该框架采用了最大更新参数化(Maximal Update Parameterization)的适配和Muon优化器,并成功应用于预训练一个155B参数的基础模型,证明了其在预测大规模MoE训练最优配置方面的有效性。 AI

影响 降低了训练大型MoE模型的计算成本,实现了更高效的开发。

排序理由 详细介绍大型模型新训练方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新框架高效预测大型MoE模型最优学习率

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
详细介绍大型模型新训练方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
20 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [1]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    循序渐进:面向大规模混合专家的计算高效超参数迁移

    A two-step hyperparameter transfer framework predicts optimal learning rates for large Mixture-of-Experts models by scaling across widths and token budgets, enabling efficient pretraining without costly sweeps.