Researchers have introduced "Muse," a novel family of optimizers designed for large language models that explores the geometric properties of parameter representations. Unlike standard Muon-style optimizers, Muse's updates are influenced by the representation of each parameter block before orthogonalization. Experiments on LLaMA2-130M and LLaMA2-600M indicate that balanced non-native representations can achieve performance comparable to native representations, while reducing the shorter dimension can weaken scaling and singular-channel support. AI
IMPACT Introduces novel optimization techniques that could improve training efficiency and performance for large language models.
RANK_REASON The cluster contains an academic paper detailing a new method for optimizers in machine learning.
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- IArxiv
- LLaMA2-130M
- LLaMA2-600M
- Marchenko--Pastur
- Muon
- Muse
- Newton-Schulz
- ScienceCast
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →