Researchers have conducted a systematic study on synthesizing high-quality pretraining data for large language models, generating over one trillion tokens. Their findings indicate that structured output formats like tables, math problems, FAQs, and tutorials are more effective than curated web text or previous synthetic methods. The study also found that generator model size beyond 1B parameters offered no additional benefit, and the selection of original data significantly impacted performance. This research led to the development of extsc{FinePhrase}, a 486-billion-token dataset that outperforms existing synthetic data baselines while reducing generation costs by up to 30 times. AI
IMPACT This research offers a scalable and cost-effective method for generating high-quality synthetic data, potentially accelerating LLM development.
RANK_REASON The cluster contains a research paper detailing a systematic study on data synthesis for LLM pretraining. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →