A new paper published on arXiv explores the impact of synthetic data on the generalization capabilities of Stochastic Gradient Descent (SGD) in high-dimensional linear regression. The research identifies that while mixed training protocols can lead to significant model collapse, a two-stage training approach, where synthetic data is used only in the initial phase, can avoid this performance floor. The study also derives scaling laws for both protocols, suggesting that larger models might exacerbate synthetic data degradation in mixed training, but high-quality pretraining with synthetic data can improve bias in two-stage methods. AI
IMPACT Highlights how training protocols critically influence the effectiveness of synthetic data, impacting model performance and scaling.
RANK_REASON Academic paper on synthetic data and SGD generalization. [lever_c_demoted from research: ic=1 ai=1.0]
- 2024
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Dohmatob et al.
- Gotit.pub
- Hugging Face
- Influence Flower
- ScienceCast
- SGD
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →