Researchers have developed two new variants of the Muon optimizer, named Muon-NSR and Muon-VS, designed to enhance the efficiency of language model pretraining. These variants adapt Muon's orthogonal momentum updates by incorporating gradient variance information, similar to Adam-style methods. Experiments on Llama-style and GPT-2 models demonstrated that Muon-VS, in particular, achieved a 1.33x speedup in reaching a target validation loss compared to the standard Muon optimizer. AI
IMPACT Introduces variance-adaptive modulation to Muon-style optimizers, potentially reducing compute costs and accelerating training for large language models.
RANK_REASON The cluster contains an academic paper detailing new methods for language model pretraining optimizers. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →