Researchers have developed LionMuon, a novel optimizer designed to reduce the significant computational cost of pretraining large language models. LionMuon alternates between computationally expensive spectral steps, similar to Muon, and cheaper sign steps, like Lion. This hybrid approach, utilizing a shared momentum buffer, aims to achieve lower loss with less training time and reduced memory footprint compared to existing optimizers such as AdamW, Lion, and Signum. Experiments on models trained on FineWeb demonstrated LionMuon's effectiveness in reaching lower losses and reducing wall-clock time. AI
IMPACT Reduces computational costs for LLM pretraining, potentially accelerating research and development.
RANK_REASON The cluster contains a research paper detailing a new method for training language models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →