stochastic gradient descent
PulseAugur coverage of stochastic gradient descent — every cluster mentioning stochastic gradient descent across labs, papers, and developer communities, ranked by signal.
- instance of ScienceCast 90%
- instance of Influence Flower 90%
- instance of Gotit.pub 70%
- instance of alphaXiv 70%
- used by Deep Neural Networks 70%
- instance of cs.LG 70%
- instance of CatalyzeX 70%
- used by Neural tangent kernel 70%
- competes with Adam optimizer 70%
- used by partial differential equations 70%
- instance of Deep Neural Networks 70%
- used by Gaussian Processes 70%
7 day(s) with sentiment data
-
New research explores advanced gradient descent for operator learning and optimization
Two new research papers explore advanced gradient descent techniques for complex optimization problems. The first paper details stochastic gradient descent (SGD) for learning operators between Hilbert spaces, establishi…
-
New method approximates stochastic gradient descent over probability measures
Researchers have developed a novel approach to approximate stochastic gradient descent (SGD) dynamics over probability measures, specifically within the Wasserstein space P2. By lifting the problem to a linear Hilbert s…
-
New Batched SGD method offers high-probability convergence guarantees
Researchers have introduced Batched SGD, a novel variant of stochastic gradient descent designed to achieve high-probability convergence guarantees for optimization problems. This method partitions online samples into e…
-
Four arXiv papers advance stochastic optimization theory · 4 sources tracked
Four new research papers published on arXiv explore advanced convergence properties of stochastic optimization methods. The first paper introduces a unified theory for steady-state convergence of stochastic approximatio…
-
Z-transform method applied to quadratic optimization in new research paper
A new paper explores the application of the z-transform method to quadratic optimization problems. The research demonstrates how this classical tool, typically used in signal processing and control theory, can yield nov…
-
New framework bounds information acquisition in neural networks
Researchers have developed a new framework to understand how neural networks acquire information during the learning process. By modeling stochastic gradient descent (SGD) as a Markovian stochastic process, they derived…
-
New research advances stochastic optimization for machine learning · 5 sources tracked
Several recent research papers explore advancements in stochastic optimization techniques, particularly focusing on gradient descent and its variants for complex machine learning problems. One paper demonstrates that va…
-
New framework analyzes attention dynamics in foundation models
Researchers have developed a new framework called attention-indexed models to better understand the training dynamics of attention mechanisms in large foundation models. This framework reveals that the optimization land…
-
Adam optimizer gets first unconditional error analysis
Researchers have developed a new theoretical framework to provide uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method. This work addresses a long-standing research pro…
-
SGD Dynamics Modeled as Percolation Process, Extending to Adam and AdamW
Researchers have modeled the dynamics of Stochastic Gradient Descent (SGD) as a percolation process, revealing how architectural symmetries cause subnetworks to merge in discrete blocks. These transitions result in vari…
-
Subspace Levenberg-Marquardt algorithms evaluated for neural network training
Researchers have evaluated subspace Levenberg-Marquardt algorithms for training neural networks, aiming to improve efficiency for larger models. These subspace methods, including Krylov subspace LM and hybrid subspace L…
-
New empirical Bayes approach enhances generalized linear models
Researchers have developed a novel empirical Bayes approach for fitting Bayesian generalized linear models, introducing a mean-field variational inference method that estimates the prior within the algorithm, making it …
-
New adaptive stopping rules boost SGD efficiency in stochastic optimization
Researchers have developed new trajectory-adaptive stopping rules for stochastic optimization algorithms like Stochastic Gradient Descent (SGD). These rules address the mismatch between theoretical fixed-time analysis a…
-
Anytime Pretraining offers horizon-free LLM training with weight averaging
Researchers have introduced "Anytime Pretraining," a novel approach to training large language models that eliminates the need for pre-defined training horizons. This method utilizes horizon-free learning-rate schedules…
-
New research explores distributed optimization with Graphon Particle Systems
Researchers have introduced Graphon Particle Systems as a method to analyze distributed optimization problems within a continuum of nodes. The study proposes stochastic gradient descent and gradient tracking algorithms …
-
New research offers theoretical convergence guarantees for neural network training
Two new research papers explore theoretical underpinnings of neural network training. The first paper establishes convergence guarantees for gradient descent in general feedforward neural networks by introducing a gener…
-
Weak correlations principle explains linearization in gradient-based learning systems
A new paper published on arXiv explores the principle of weak correlations as the underlying reason for the linearization observed in gradient-based learning systems. The research suggests that the simplified dynamics s…
-
Researchers Analyze Stochastic Gradient Descent with Discontinuities
Researchers have analyzed stochastic gradient descent (SGD) when applied to loss functions that exhibit discontinuity across lower-dimensional manifolds. The study focuses on the differential equation limit of SGD to un…
-
New research outlines SGD preconditioner design for stability and noise reduction
A new research paper published on arXiv details design criteria for stochastic gradient descent (SGD) preconditioners, focusing on local conditioning, noise floors, and basin stability. The paper derives bounds where co…
-
Research paper details "benign misfitting" in linear regression models
A new research paper explores the phenomenon of "benign misfitting" in linear regression models, where a model that performs poorly on training data can still generalize well to new, unseen data. This occurs in a specif…