PulseAugur
中
实时 18:50:22
English(EN) Gradient Descent as Implicit EM in Distance-Based Neural Models

新论文将梯度下降与神经网络中的隐式EM联系起来

一篇新论文提出,在某些神经网络目标函数中,梯度下降表现为一种隐式的期望最大化(EM)算法。研究表明,对于涉及距离或能量的log-sum-exp结构的优化目标,相对于每个距离的梯度恰好是相应分量的负后验责任。这个代数恒等式是Fisher恒等式的一个特例,意味着标准的神经网络训练在没有显式辅助变量的情况下,隐式地执行了广义EM。这些发现统一了无监督混合模型、注意力机制和交叉熵分类,解释了Transformer等模型中观察到的软聚类和贝叶斯不确定性跟踪等现象。 AI

影响 提供了一个理论框架,可能带来更高效、更易于解释的神经网络训练。

排序理由 该集群包含一篇学术论文,详细介绍了神经网络训练动力学的新理论见解。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新论文将梯度下降与神经网络中的隐式EM联系起来

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇学术论文,详细介绍了神经网络训练动力学的新理论见解。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
93 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Alan Oursland ·

    基于距离的神经网络模型中的梯度下降作为隐式EM

    arXiv:2512.24780v2 Announce Type: replace Abstract: Neural networks trained with standard objectives exhibit behaviors characteristic of probabilistic inference: soft clustering, prototype specialization, and Bayesian uncertainty tracking. These phenomena appear across architectu…