Researchers have introduced WSqD, a novel learning rate schedule designed for large model training that is independent of the training horizon. Unlike existing methods like cosine annealing and Warmup-stable-decay (WSD), WSqD's base schedule does not require a predetermined training duration, allowing for flexible extension of training. This approach, inspired by stochastic convex optimization, theoretically achieves optimal convergence rates and has been empirically shown to match or exceed baseline performance on language model pretraining tasks using the SlimPajama corpus. AI
IMPACT This horizon-free learning rate schedule could simplify and improve the efficiency of training large language models.
RANK_REASON The cluster contains an academic paper detailing a new method for large model training.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →