PulseAugur
实时 15:56:18
English(EN) Scaling Limits of Constant-Stepsize SGD at Flat Minima

新研究详解SGD在平坦极小值下的缩放极限

一篇新论文探讨了随机梯度下降(SGD)应用于具有平坦极小值的凸目标函数时的缩放极限。研究表明,对于这类目标函数,SGD的行为会发生根本性变化,导致不同的缩放定律,并且当步长接近于零时可能出现非高斯极限。这些发现对于理解损失曲率不是强凸情况下的优化场景尤为重要。 AI

影响 为与机器学习模型训练相关的优化算法提供了理论见解。

排序理由 该集群包含一篇详细介绍优化算法理论研究的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv stat.ML 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新研究详解SGD在平坦极小值下的缩放极限

报道来源 [1]

  1. arXiv stat.ML TIER_1 English(EN) · Jingyi Zhang, Cheng Mao, Debankur Mukherjee ·

    平坦极小值下恒定步长SGD的缩放极限

    arXiv:2607.16384v1 Announce Type: cross Abstract: For stochastic gradient descent (SGD) with a constant stepsize $\alpha$, the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons. In the strongly convex case, thi…