Sharpness aware minimization
PulseAugur coverage of Sharpness aware minimization — every cluster mentioning Sharpness aware minimization across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New analysis quantifies SAM's bias toward flat minima
Researchers have analyzed the implicit bias of Sharpness-Aware Minimization (SAM) in improving model generalization. Their linear stability analysis reveals a quantitative relationship between SAM's perturbation radius …
-
Muon optimizer shows promise in theoretical and practical neural network training
Two new research papers explore the Muon optimizer, an approach designed to better handle matrix-structured parameters in neural networks. The first paper introduces a matrix-aware geometry for Sharpness-Aware Minimizat…
-
New GEAR-SAM method enhances model generalization and robustness
Researchers have developed Gradient-Energy Adaptive Radius SAM (GEAR-SAM), a novel approach to Sharpness-Aware Minimization (SAM) that aims to improve model generalization and robustness. Unlike standard SAM, which allo…
-
New EISAM optimizer enhances deep learning generalization
Researchers have introduced Extragradient-Inspired Sharpness-Aware Minimization (EISAM), a new optimizer designed to improve generalization in deep learning. EISAM employs a two-step process, involving a prediction and …
-
Federated Learning Research Tackles Client Drift via Frequency Domain Analysis · 2 sources tracked
Two new research papers, FedFFT and SpecGradFilter, propose novel methods to address client drift in federated learning by analyzing gradient perturbations in the frequency domain. Both papers identify that inconsistenc…
-
New research shows Sharpness-Aware Minimization improves AI model calibration
A new research paper explores how Sharpness-Aware Minimization (SAM) can improve the calibration of deep neural networks, making them less prone to overconfidence in critical applications. The study suggests SAM implici…
-
New TALAS framework improves language model distillation efficiency
Researchers have introduced TALAS, a novel framework for knowledge distillation in pre-trained language models. TALAS synchronizes hierarchical alignment with advanced optimization techniques to improve efficiency and p…
-
New research probes SAM optimizer's stability and adaptive learning
Two new research papers delve into the complexities of Sharpness-Aware Minimization (SAM), a popular deep learning training technique. The first paper analyzes SAM's convergence instability near saddle points, theoretic…
-
New method improves deep learning generalization with unlabeled data
Researchers have developed a new method called Inconsistency-Aware Minimization (IAM) to improve how deep learning models generalize, particularly when using unlabeled data. IAM introduces a novel measure called local i…
-
New LE-SAM method boosts model generalization over traditional SAM
Researchers have introduced Loss-Equated SAM (LE-SAM), a novel approach to enhance generalization in machine learning models. This method addresses a mismatch in Sharpness-Aware Minimization (SAM) by focusing on a fixed…
-
New methods like SMF and SAM reduce catastrophic forgetting in LLMs
Two new research papers explore methods to mitigate catastrophic forgetting in language models during fine-tuning. One paper introduces Sparse Memory Finetuning (SMF), which adds memory layers and updates only heavily a…