PulseAugur
实时 07:47:12
English(EN) The Loss Does Not See the Basis, But Adam Does [R]

Adam优化器打破了因子模型中的低秩偏差,而梯度下降则不会

一篇新论文揭示,虽然梯度下降由于规范等变性而隐式地偏好因子矩阵模型中的低秩解,但流行的Adam优化器却不会。这种差异归因于Adam的每坐标二阶矩计算,它打破了梯度下降所遵循的对称性。实验表明,像Adam和RMSProp这样的优化器会丢失这种低秩偏差,导致在矩阵感知和Transformer等任务上的性能不如梯度下降和其他等变优化器。研究表明,各向异性而不是自适应性是影响这种偏差的关键因素。 AI

影响 这项研究强调了Adam和梯度下降等优化器在处理低秩结构方面的根本区别,这可能会影响各种AI任务的模型训练和性能。

排序理由 该集群包含一篇详细介绍优化器行为新发现的研究论文。

在 r/MachineLearning 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

Adam优化器打破了因子模型中的低秩偏差,而梯度下降则不会

报道来源 [2]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    损失不看基础,但 Adam 看

    Optimizer behavior in factored matrix models depends on gauge equivariance, with coordinate-wise methods breaking low-rank bias and causing divergent solutions in transformers and sensing tasks.

  2. r/MachineLearning TIER_1 English(EN) · /u/EtherealGlyph ·

    损失不看基础,但Adam看[R]

    <table> <tr><td> <a href="https://www.reddit.com/r/MachineLearning/comments/1vmjb3p/the_loss_does_not_see_the_basis_but_adam_does_r/"> <img alt="The Loss Does Not See the Basis, But Adam Does [R]" src="https://preview.redd.it/cldvfu1oyyih1.png?width=140&amp;height=54&amp;auto=web…