PulseAugur
EN
LIVE 09:01:00

New framework efficiently predicts optimal learning rates for large MoE models

Researchers have developed a novel two-step framework to efficiently determine optimal learning rates for large Mixture-of-Experts (MoE) models. This method leverages hyperparameter transfer across different model widths and extrapolates findings to massive token budgets, significantly reducing the computational cost associated with traditional hyperparameter sweeps. The framework, which uses a Maximal Update Parameterization adaptation and the Muon optimizer, was successfully applied to pretrain a 155B-parameter foundation model, demonstrating its effectiveness in predicting optimal configurations for large-scale MoE training. AI

IMPACT Reduces computational costs for training large MoE models, enabling more efficient development.

RANK_REASON Academic paper detailing a new method for training large models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New framework efficiently predicts optimal learning rates for large MoE models

COVERAGE [1]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

    A two-step hyperparameter transfer framework predicts optimal learning rates for large Mixture-of-Experts models by scaling across widths and token budgets, enabling efficient pretraining without costly sweeps.