A new research paper analyzes the pre-training dynamics of language models from the perspective of local landscape geometry. The study identifies two distinct phases: Phase I, where sharpness leads to instability with large learning rates and necessitates learning rate warmup, and Phase II, where gradient noise scale governs the landscape. The research proposes a dynamic batch-size scheduler that increases batch size late in training, offering actionable strategies for optimizing large-scale pre-training. AI
IMPACT Offers new insights into optimizing language model pre-training efficiency and stability.
RANK_REASON The cluster contains a research paper detailing new findings and proposed methods for language model pre-training. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →