PulseAugur
实时 05:05:44

新研究将优化器选择与LLM微调中遗忘减少联系起来

研究人员探讨了优化器一致性对大型语言模型微调的影响。一项研究表明,在预训练和微调过程中使用相同的优化器可以减少知识遗忘,并在新任务上获得更好的性能,这种现象被称为“优化器-模型一致性”。与LoRA等其他方法相比,这种方法可能提供更好的学习-遗忘权衡。另一篇论文引入了“谱边分析”来研究神经网络训练中的相变,将“grokking”和能力提升等现象与参数更新矩阵的谱隙联系起来。该框架表明,优化器的选择会影响这些动态,实验结果证实了在各种模型尺寸上的预测。 AI

影响 这些研究为理解和改进大型语言模型的训练和微调提供了新的理论框架和经验证据,有望带来更高效、更有效的模型开发。

排序理由 两篇在arXiv上发表的学术论文,详细介绍了神经网络训练动态和优化方面的新发现。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新研究将优化器选择与LLM微调中遗忘减少联系起来

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
两篇在arXiv上发表的学术论文,详细介绍了神经网络训练动态和优化方面的新发现。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
112 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [3]

  1. arXiv cs.LG TIER_1 English(EN) · Yuxing Liu, Jianyu Wang, Tong Zhang ·

    优化器-模型一致性:使用与预训练相同的优化器进行完全微调可减少遗忘

    arXiv:2605.06654v1 Announce Type: new Abstract: Optimizers play an important role in both pretraining and finetuning stages when training large language models (LLMs). In this paper, we present an observation that full finetuning with the same optimizer as in pretraining achieves…

  2. arXiv cs.LG TIER_1 English(EN) · Yongzhong Xu ·

    光谱边缘动力学:神经网络训练中相变的一个解析-经验研究

    arXiv:2603.28964v3 Announce Type: replace Abstract: We develop the spectral edge analysis: phase transitions in neural network training -- grokking, capability gains, loss plateaus -- are controlled by the spectral gap of the rolling-window Gram matrix of parameter updates. In th…

  3. arXiv cs.AI TIER_1 English(EN) · Tong Zhang ·

    优化器-模型一致性:使用与预训练相同的优化器进行完全微调可减少遗忘

    Optimizers play an important role in both pretraining and finetuning stages when training large language models (LLMs). In this paper, we present an observation that full finetuning with the same optimizer as in pretraining achieves a better learning-forgetting tradeoff, i.e., fo…