Researchers have developed SkewAdam, a novel optimizer designed to significantly reduce memory usage during the training of Mixture-of-Experts (MoE) models. By allocating different levels of precision for optimizer state across the MoE's distinct parameter populations (backbone, experts, and router), SkewAdam drastically cuts memory requirements from 50.6 GB to 1.29 GB. This optimization allows for the training of large MoE models on standard 40 GB accelerators without compromising accuracy, as demonstrated by SkewAdam achieving better validation perplexity than AdamW, Muon, and Lion. AI
IMPACT Enables training of larger MoE models on existing hardware, potentially accelerating research and deployment.
RANK_REASON The cluster describes a new research paper detailing a novel optimization technique for training large AI models.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →