Researchers have identified a mechanism for delayed generalization, known as grokking, in linear models trained with specific optimization techniques and weight decay. Their analysis reveals a 'grokking subspace' within the empirical null space where training predictions remain stable, allowing weight decay to drive a slow relaxation process. This theory predicts grokking time based on factors like regularization strength and optimizer choice, and has been validated in synthetic models and modular addition tasks. AI
IMPACT Provides a theoretical framework for understanding and potentially controlling delayed generalization in AI models.
RANK_REASON The cluster contains a research paper detailing a new theoretical framework for understanding a phenomenon in machine learning. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries
- Hugging Face
- Influence Flower
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →