Hessian
PulseAugur coverage of Hessian — every cluster mentioning Hessian across labs, papers, and developer communities, ranked by signal.
7 day(s) with sentiment data
-
New theoretical bounds improve Langevin sampling for complex distributions
Researchers have developed new theoretical bounds for the Moreau--Yosida unadjusted Langevin algorithm (MYULA), a method used for sampling from complex probability distributions. The study focuses on nonsmooth composite…
-
New theory explains curvature in Probabilistic Circuits
Researchers have developed a compositional theory for understanding curvature in Probabilistic Circuits (PCs), a type of generative model. They demonstrated that the Hessian trace, a measure of loss-surface curvature, c…
-
Generalized Quadratic Gradient framework unifies optimization methods
Researchers have introduced Generalized Quadratic Gradient (GQG), a novel optimization framework that unifies and extends existing second-order optimization methods. GQG abstracts the core principles of Quadratic Gradie…
-
New research links mini-batch noise to loss landscape sharpness in SGD
A new research paper proposes that mini-batch noise during Stochastic Gradient Descent (SGD) training influences the sharpness of the loss landscape by causing fluctuations within the dominant subspace. The authors argu…
-
New kinetic theory formalizes zeroth-order Newton methods
Researchers have developed a formal kinetic theory for zeroth-order Newton-type methods, which are useful when gradients and Hessians are unavailable. The framework includes a Gaussian-Stein correction to accurately est…
-
Quadratic Model Shows Predictive Power for LLM Optimization Dynamics
A new paper published on arXiv proposes that the simple quadratic model can be a surprisingly accurate predictor of optimization dynamics in large language models (LLMs). Researchers demonstrated that by analyzing the H…
-
New C-PTQ method enhances multimodal LLM quantization efficiency
Researchers have developed C-PTQ, a novel post-training quantization method designed to improve the efficiency of multimodal large language models (MLLMs). This technique addresses performance degradation caused by outl…
-
Neural network research reveals functional equivalence and geometric diversity
A new research paper explores the concept of functional equivalence in neural networks, building upon the Universal Approximation Theorem. The study reveals that multiple neural network configurations can achieve identi…
-
New Gaussian Invariant MCMC methods boost statistical efficiency
Researchers have developed novel sampling methods, including Gaussian invariant versions of Random Walk Metropolis (RWM), Metropolis-adjusted Langevin algorithm (MALA), and a second-order Hessian or Manifold MALA. These…
-
Neural network Hessian eigenvalues explained by approximate symmetries
Researchers have developed a new explanation for the numerous near-zero eigenvalues of the Hessian matrix in neural networks. They propose that these vanishing eigenvalues stem from approximate symmetries within the net…
-
Gradient Descent Theory Extended for Complex Minima and Vector Outputs
This paper extends the theory of gradient descent (GD) with large step sizes to more complex scenarios. It addresses overparameterized least-squares problems with vector-valued outputs and analyzes neighborhoods of mani…
-
New research details Zeroth-Order optimization stability in deep learning
A new research paper explores the stability dynamics of Zeroth-Order (ZO) optimization methods, particularly in the context of deep learning. The study identifies a specific step size condition that governs the linear s…
-
Genetic algorithms mimic clipped gradient descent in high-dimensional AI search
Researchers have demonstrated that genetic algorithms can effectively function as a form of clipped gradient descent in high-dimensional search spaces. This process involves mutation-selection mechanisms that implicitly…
-
New CWGD method improves optimization noise measurement for deep learning
Researchers have developed a new method called Curvature-Weighted Gradient Diversity (CWGD) to better measure optimization noise in deep learning models. Unlike traditional methods that treat all parameter directions eq…
-
Hessian Eigenvector Dynamics Reveal Optimizer Differences in Neural Network Training
Researchers have analyzed the evolution of Hessian eigenvectors during neural network training, revealing distinct behaviors between different optimizers. The study found that SGD tends to stabilize leading curvature di…
-
New method offers second-order KKT guarantees for Bregman ADMM
Researchers have developed a novel approach to analyze Bregman ADMM for nonconvex and non-Lipschitz optimization problems. This method replaces the standard Lipschitz gradient assumption with a two-sided relative smooth…
-
New research frameworks model gradient descent at the edge of stability
Two new research papers explore the phenomenon of gradient descent operating at the edge of stability (EoS) in deep learning. The first paper introduces 'Edge Flow,' a system of differential equations that models gradie…
-
New 'Architecture Warm-Up' Stabilizes Transformer Training
Researchers have developed a new method to stabilize the training of large Transformer models, which are often prone to instability and divergence. The approach, called "architecture warm-up," involves progressively inc…
-
New optimization techniques emerge for faster, more efficient AI model training · 8 sources tracked
Several recent arXiv papers explore advancements in optimization techniques for machine learning. Researchers have proposed new methods like Weight Adaptation ASNG (WA-ASNG) to improve parallel performance in evolutiona…
-
New research explores how network symmetry aids optimization in overparameterized deep learning models.
A new paper analyzes how overparameterization in neural networks aids optimization by introducing additional symmetries. These symmetries act as a form of preconditioning on the Hessian, leading to better-conditioned mi…