Researchers have developed a compute-efficient framework for determining optimal learning rates in large Mixture-of-Experts (MoE) models. This two-step process involves transferring optimal learning rates across models of varying widths and then extrapolating to massive token budgets. By using linear regression on small proxy models, the method accurately predicts ideal learning rates for training horizons up to 10 trillion tokens, significantly reducing computational costs. The framework was successfully applied to pretrain a 155B parameter foundation model, validating its effectiveness. AI
IMPACT Reduces computational costs for training large MoE models, enabling more efficient scaling.
RANK_REASON Academic paper detailing a new methodology for training large models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →