A new paper reveals that while gradient descent implicitly favors low-rank solutions in factored matrix models due to gauge equivariance, the popular Adam optimizer does not. This difference is attributed to Adam's per-coordinate second moment calculation, which breaks the symmetry that gradient descent respects. Experiments show that optimizers like Adam and RMSProp lose this low-rank bias, leading to poorer performance on tasks like matrix sensing and transformers compared to gradient descent and other equivariant optimizers. The research suggests that anisotropy, rather than adaptivity, is the key factor affecting this bias. AI
IMPACT This research highlights a fundamental difference in how optimizers like Adam and gradient descent handle low-rank structures, potentially impacting model training and performance on various AI tasks.
RANK_REASON The cluster contains a research paper detailing novel findings about optimizer behavior.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →