A new research paper titled "Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack" challenges the common assumption that the best pretraining checkpoint will yield the best results after subsequent training. The study, conducted on a 30 billion parameter mixture-of-experts model, found that checkpoints performing better after the full downstream training stack exhibited higher solution density. This means these checkpoints retained downstream performance even when subjected to local weight perturbations, suggesting a more robust starting point for further development. AI
IMPACT Findings suggest that careful selection of pretraining checkpoints is crucial for optimizing downstream model performance and robustness.
RANK_REASON The cluster contains a research paper detailing findings on language model training checkpoints. [lever_c_demoted from research: ic=1 ai=1.0]
- 30b Parameter Model
- arXiv
- Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack
- Hugging Face
- mixture of experts
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →