Researchers have identified a new critical factor in deep neural network training instability, termed 'weight-norm criticality.' This phenomenon, distinct from the commonly understood 'learning-rate criticality,' arises from the interplay between normalization techniques and weight decay. As weight decay increases, it can drive parameter norms toward zero, leading to a sharper loss landscape and abrupt loss spikes that destabilize optimization dynamics. This finding offers a mechanistic explanation for why excessive weight decay, while potentially improving generalization, can ultimately hinder training. AI
IMPACT Provides a new theoretical framework for understanding and potentially mitigating training instabilities in deep learning models.
RANK_REASON The cluster contains two identical arXiv papers detailing a new theoretical mechanism for understanding training instability in deep neural networks.
- data normalization
- deep neural network
- Edge of Stability
- generalization
- learning-rate criticality
- Loss Landscape
- loss spikes
- Tikhonov regularization
- Weight-norm Criticality
- arXiv
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →