Researchers have developed a statistical theory to explain the phenomenon of "grokking," where a model learns underlying signals at a different stage than it fits training data. The theory characterizes how regularization geometry and signal sparsity influence generalization near interpolation, particularly in high-dimensional regression. Experiments on diagonal linear networks and transformers demonstrate the theory's predictions, revealing a statistical instability in minimum-norm interpolation where small changes in regularization strength can lead to significantly different generalization while maintaining low training error. AI
IMPACT Provides theoretical grounding for understanding model generalization, potentially informing future model design and training strategies.
RANK_REASON The item is an academic paper published on arXiv detailing a new statistical theory for a machine learning phenomenon. [lever_c_demoted from research: ic=1 ai=1.0]
- all-zero predictor
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- convex norms
- DagsHub
- Diagonal Linear Networks
- Gotit.pub
- grokking
- Hugging Face
- machine learning
- modular arithmetic
- regression analysis
- ScienceCast
- transformers
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →