Researchers have developed a novel two-step framework to efficiently determine optimal learning rates for large Mixture-of-Experts (MoE) models. This method leverages hyperparameter transfer across different model widths and extrapolates findings to massive token budgets, significantly reducing the computational cost associated with traditional hyperparameter sweeps. The framework, which uses a Maximal Update Parameterization adaptation and the Muon optimizer, was successfully applied to pretrain a 155B-parameter foundation model, demonstrating its effectiveness in predicting optimal configurations for large-scale MoE training. AI
IMPACT Reduces computational costs for training large MoE models, enabling more efficient development.
RANK_REASON Academic paper detailing a new method for training large models. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →