PulseAugur
中
实时 08:53:58
English(EN) Do Flatter Minima Drive Better Generalization? An Algorithmic Separation in Grokking

研究发现,仅靠平坦度不足以实现神经网络的 grokking

研究人员调查了损失景观平坦度在神经网络泛化中的作用,特别是在 grokking 现象中。虽然之前的工作表明平坦度是泛化所必需的,但本研究发现,仅使用 Sharpness Aware Minimization (SAM)(它倾向于将训练导向更平坦的解)不足以可靠地诱导 grokking。然而,当 SAM 与权重衰减结合使用时,它能将泛化解的过渡速度在 epoch 层面提高多达四倍。对最小的两层 ReLU 模型的理论分析表明,仅靠平坦度无法区分记忆解和泛化解,但权重衰减有利于泛化,而 SAM 可以通过破坏记忆插值器来加速这一过渡。 AI

影响 提供了对训练动态如何影响模型泛化更细致的理解,可能指导未来的优化技术。

排序理由 学术论文,详细介绍了关于神经网络泛化的新发现。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究发现,仅靠平坦度不足以实现神经网络的 grokking

本文如何被排名

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
学术论文,详细介绍了关于神经网络泛化的新发现。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Mohnish Harwani ·

    更平坦的最小值是否能带来更好的泛化能力?Grokking算法中的分离分析

    arXiv:2610.11206v1 Announce Type: new Abstract: Flat loss landscapes have long been linked to better generalization in neural networks. However, its role as a causal mechanism for generalization is less established. Grokking provides an unique testbed to understand this distincti…