Researchers have identified a limitation in the Muon optimization method, which is designed for matrix-valued parameters in machine learning. While Muon orthogonalizes momentum matrices, it can fail to converge to a global optimum for convex Lipschitz objectives, even with adaptive step sizes. The study proposes MAGD, an alternative that combines orthogonalized momentum with gradient weighting, offering improved convergence rates and practical performance in experiments including LLM pretraining. AI
IMPACT Introduces a more reliable optimization method that could improve training efficiency for large language models.
RANK_REASON Academic paper introducing a new optimization method with theoretical analysis and experimental results. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →