muon
PulseAugur coverage of muon — every cluster mentioning muon across labs, papers, and developer communities, ranked by signal.
12 day(s) with sentiment data
Aurora optimizer may outperform Muown in addressing Muon's neuron death
Tilde Research's Aurora optimizer is specifically designed to fix 'neuron death' in Muon, a problem not explicitly addressed by Muown. While Muown improves spectral norm drift, Aurora's targeted approach to neuron inactivity could lead to more comprehensive performance gains, especially in scenarios where neuron death is a primary bottleneck.
Muon optimizer's spectral norm drift is a key area for improvement
Multiple recent papers (Muown, Pion, and the general mode connectivity research) highlight issues related to spectral norms and Muon. Muown explicitly addresses 'upward drift of spectral norms', while Pion aims to 'preserve spectrum'. This suggests that managing spectral properties is a critical challenge for Muon's stability and performance.
Spectrum preservation is a common theme in new optimizer research
The introduction of Pion, which 'preserves spectrum', and Muown, which addresses 'spectral norm drift', indicates a broader trend in optimizer development. This focus on maintaining spectral properties suggests that current optimizers, including Muon, may suffer from spectral instability that hinders training.
Muon's spectral properties are being actively studied in relation to optimizer behavior and mode connectivity
Multiple recent clusters highlight research into Muon's spectral properties and how they interact with optimization dynamics. The connection between optimizers, spectral norms, and mode connectivity suggests ongoing theoretical and empirical work is exploring fundamental aspects of Muon's behavior.
Muon's neuron death issue may be addressed by new optimizers like Aurora within 3 months
The Tilde Research launch of Aurora specifically targets neuron death in Muon. Given Aurora's public release and demonstrated effectiveness, it's plausible that Muon users will adopt Aurora or similar solutions to mitigate this issue within the next quarter.
-
New derivative-free framework adapts Muon for optimization without gradients
Researchers have developed a derivative-free framework to adapt the Muon optimization method for scenarios where gradients are unavailable or unreliable. This new approach uses structured finite differences to construct…
-
Temperon training method achieves SAM quality with reduced cost
Researchers have introduced Temperon, a novel training method designed to achieve the quality of Sharpness-Aware Minimization (SAM) while significantly reducing computational costs. Temperon utilizes a two-phase approac…
-
New method offers certified early stopping for AI regularized inverse problems
Researchers have developed a method for certified early stopping in regularized inverse problems, which involves trading off a data-fidelity term against a regularizer. This approach utilizes an exact duality-gap identi…
-
New optimization method MAGD shows promise for LLM pretraining
Researchers have identified a limitation in the Muon optimization method, which is designed for matrix-valued parameters in machine learning. While Muon orthogonalizes momentum matrices, it can fail to converge to a glo…
-
7B model ZGCM-1 prioritizes tool use and large context over memorization
Researchers from Zhongguancun Academy and Zhongguancun Institute of AI have developed ZGCM-1, a 7.39B parameter model that prioritizes tool use and a large context window over memorizing vast datasets. This approach all…
-
Deeper, thinner models outperform wider ones in sub-150M parameter regime
Researchers have explored the impact of model depth versus width in the sub-150 million parameter range, finding that a deeper, thinner architecture (23 layers x 576 hidden) outperformed a wider, shallower one (53.5M vs…
-
7 PhD students train 7B LLM from scratch using hundreds of AI agents
Seven doctoral students from Beijing Zhongguancun Academy successfully trained a 7B large language model, ZGCM-1, from scratch in just three months. They achieved this by leveraging a team of hundreds of AI agents to ha…
-
New ISO-LoRA optimizer boosts parameter-efficient adaptation for LLMs
Researchers have introduced ISO-LoRA, a novel optimization technique designed to enhance the efficiency of Low-Rank Adaptation (LoRA) for large language models. Unlike traditional LoRA methods that focus solely on the r…
-
New Musec Optimizer Enhances LLM Training Stability
Researchers have introduced MomentUm SpEctral Clipping (Musec), a novel optimizer designed to stabilize the training of large language models. Musec addresses instability issues inherent in the Muon optimizer, which oft…
-
Open-source ZGCM-1 model achieves high efficiency in math and agentic search
Researchers have introduced ZGCM-1, a 7B parameter foundation model designed for mathematical reasoning and agentic search. The model leverages an efficient training recipe that combines architectural innovations like i…
-
New research explains why optimizers struggle with equivariant networks
Researchers have identified a key reason why certain optimizers like Muon outperform Adam when training equivariant neural networks. The issue stems from how Adam handles learning rates across different blocks within an…
-
New Quadratic Spectral Descent method improves GPT pre-training efficiency
Researchers have developed a new optimization method called Quadratic Spectral Descent (QSD) that improves upon the existing Muon algorithm for training large language models. QSD incorporates local curvature informatio…
-
New FedSubMuon method slashes LLM federated fine-tuning communication costs
Researchers have developed FedSubMuon, a novel method for federated fine-tuning of large language models (LLMs) that significantly reduces communication costs. This approach optimizes compact coefficient matrices within…
-
Muon-C optimizer achieves superior performance on convolutional kernels
Researchers have introduced Muon-C, a novel operator-aligned optimizer designed for convolutional kernels. This new method represents kernel momentum as frequency-wise channel-transfer matrices, which are then independe…
-
New research shows optimizer performance shifts with training horizon
A new research paper explores how different optimizers perform as training horizons and parameter counts increase. The study found that the optimal hyperparameters and relative performance of optimizers like Muon, SOAP,…
-
LLaDA-Image sets new open-source SOTA for image generation
Researchers have introduced LLaDA-Image, a novel framework for generating high-quality images using a 6B Diffusion Transformer trained from scratch. This model leverages image-only pre-training and a specialized optimiz…
-
New LoRA-TSD optimizer offers cheaper, faster fine-tuning for LLMs
Researchers have developed LoRA-TSD, a novel optimizer for fine-tuning large language models. This method treats each update as a tangent vector on a fixed-rank matrix manifold, employing a spectral-norm steepest-descen…
-
New Muon optimizer variants boost language model pretraining efficiency
Researchers have developed two new variants of the Muon optimizer, named Muon-NSR and Muon-VS, designed to enhance the efficiency of language model pretraining. These variants adapt Muon's orthogonal momentum updates by…
-
Qwen releases Qwen3.8-Flash-Next multimodal MoE model with 1M context
The Qwen team has released the weights for their Qwen3.8-Flash-Next model, a multimodal Mixture-of-Experts (MoE) architecture. This new model incorporates innovations such as Gated DeltaNet+Qwen Sparse Attention (GDN+QS…
-
OrScale optimization method enhances neural network training
Researchers have introduced OrScale, a novel optimization method designed to improve the training of large neural networks. OrScale addresses the direction and magnitude of updates by adapting the trust-ratio principle …