Researchers have developed a unified framework using Bregman divergences to analyze the impact of different divergence choices on neural network optimizers, specifically Shampoo. This framework connects various popular divergences, allowing for a joint study of their effects. The study empirically analyzes how divergence selection influences Kronecker approximation and interacts with finite-sample errors in preconditioning, suggesting that certain divergences can better mitigate underestimation of the empirical second moment. These findings were validated through GPT-2 pretraining experiments, offering guidance for improving Shampoo and related optimization techniques. AI
IMPACT Provides a theoretical framework to guide the development of more effective neural network optimizers, potentially leading to faster and more stable training.
RANK_REASON The cluster contains an academic paper detailing a new theoretical framework and empirical validation for improving neural network optimizers. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Bregman divergence
- Frobenius divergence
- GPT-2
- Hugging Face
- Kronecker approximation
- Kullback--Leibler divergence
- Shampoo
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →