A new optimizer called SkewAdam has been developed to significantly reduce the memory required for training Mixture-of-Experts (MoE) models. This optimizer achieves a 97.4% reduction in optimizer state memory by employing a tiered allocation strategy, treating backbone parameters, expert parameters, and router parameters with different levels of precision. This optimization allows a 6.78 billion parameter MoE model to fit on a single 40GB GPU without compromising convergence or stability. AI
IMPACT Enables training of larger MoE models on more accessible hardware, potentially accelerating research and development in this area.
RANK_REASON The item describes a new optimizer published in a preprint, detailing its technical approach and performance improvements. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →