Two new arXiv papers explore the phenomenon of 'grokking' in neural networks, where models generalize only after memorizing training data. One paper proposes 'Low-Rank Decay' (LRD) as a spectral regularizer to improve grokking by compressing singular values, showing it can accelerate rank collapse and expand the data-fraction boundary for generalization. The other paper frames grokking as a constrained optimization problem, demonstrating that gradient descent minimizes weight norm on the zero-loss manifold and deriving a closed-form expression for post-memorization dynamics. AI
IMPACT These papers offer theoretical insights into delayed generalization in neural networks, potentially guiding future model training strategies.
RANK_REASON Two academic papers published on arXiv discussing a specific phenomenon in neural networks.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →