Researchers have developed new hyperparameter scaling laws specifically for ultra-sparse Mixture-of-Experts (MoE) models, addressing challenges in transferring optimal learning rates and batch sizes across varying sparsity levels. Through extensive pre-training runs involving 1,800 experiments and approximately 20 trillion tokens, they identified two distinct scaling regimes. The findings indicate that optimal batch size scales with training tokens, while learning rate is influenced by training compute and robust to model size and data allocation. These new laws incorporate activation ratio as a multiplicative factor, enabling more accurate hyperparameter prediction for sparse MoEs, even for models with billions of parameters and very low expert activation. AI
IMPACT Enables more efficient training and deployment of sparse MoE models by providing accurate hyperparameter guidance.
RANK_REASON Academic paper detailing new scaling laws for MoE models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- IArxiv Recommender
- mixture of experts
- Nvidia H800
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →