A new research paper explores the differing behaviors of optimization algorithms like Adam and gradient descent when applied to factored models. The study reveals that while gradient descent is implicitly biased towards low-rank solutions due to the loss function's gauge symmetry, Adam and similar coordinate-wise optimizers do not share this bias. This difference is attributed to the gauge equivariance property, which is necessary for transferring properties from gradient flow but not sufficient for low-rank recovery. The research sorts nine update rules by recovery error, finding that Adam separates gauge-equivalent initializations in transformers, leading to significant differences in per-head invariants. AI
IMPACT Explains fundamental differences in how optimizers like Adam and gradient descent behave with factored models, impacting model training and recovery.
RANK_REASON Research paper detailing novel findings about optimization algorithms. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →