Researchers have developed an open-source pretraining recipe that significantly reduces the cost of training large language models, making them accessible on consumer GPUs for under $7,000. Their Puro-2B model, trained on RTX 5090 GPUs, achieves performance comparable to larger models like Qwen2-1.5B. The study also introduces a "Puro Cost Scaling Law" which estimates that reaching Qwen2-1.5B performance costs less than $5,090, and examines how data curricula influence downstream performance. AI
IMPACT Lowers the barrier for academic and open-source communities to train large language models on accessible hardware.
RANK_REASON The item describes a new open-source training recipe and a model released in a research paper. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →