Researchers have investigated the role of flat loss landscapes in neural network generalization, particularly within the phenomenon of grokking. While previous work suggested flatness is necessary for generalization, this study found that using Sharpness Aware Minimization (SAM) alone, which biases training towards flatter solutions, was insufficient to reliably induce grokking. However, when SAM was combined with weight decay, it accelerated the transition to generalizing solutions by up to four times at the epoch level. Theoretical analysis on a minimal two-layer ReLU model indicated that flatness alone cannot distinguish memorizing from generalizing solutions, but weight decay favors generalization, and SAM can accelerate this transition by destabilizing memorizing interpolants. AI
IMPACT Provides a more nuanced understanding of how training dynamics influence model generalization, potentially guiding future optimization techniques.
RANK_REASON Academic paper detailing novel findings on neural network generalization. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- gradient descent
- grokking
- Hugging Face
- Neural Networks
- Sam
- Sharpness aware minimization
- Tikhonov regularization
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →