PulseAugur
实时 14:40:41
English(EN) Scale Weight Decay and Train Better

新的缩放权重衰减方法加速神经网络训练

研究人员引入了一种新颖的缩放权重衰减方法,该方法受Robbins-Monro条件的启发,以改进神经网络训练。该技术根据峰值学习率的比例调整权重衰减,确保渐近平稳性,并避免恒定权重衰减引入的偏差。当应用于使用Muon优化器的混合专家模型时,这种缩放权重衰减(Muon-SW)显示出显著的加速效果,在大规模训练中,达到目标验证损失的速度最多可提高30%。 AI

影响 有望以最小的实施工作量显著加速前沿模型的预训练。

排序理由 该集群包含一篇详细介绍神经网络新训练方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的缩放权重衰减方法加速神经网络训练

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Anuj Apte ·

    Scale Weight Decay and Train Better

    arXiv:2607.23777v1 Announce Type: cross Abstract: The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data. This is typically done with a constant decoupled weight decay which causes the network weights to shrink steadily over the…