PulseAugur
EN
LIVE 11:50:07
ENTITY Hessian

Hessian

PulseAugur coverage of Hessian — every cluster mentioning Hessian across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
9
20 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
9
20 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

7 day(s) with sentiment data

RECENT · PAGE 1/1 · 20 TOTAL
  1. TOOL · CL_200182 ·

    New theoretical bounds improve Langevin sampling for complex distributions

    Researchers have developed new theoretical bounds for the Moreau--Yosida unadjusted Langevin algorithm (MYULA), a method used for sampling from complex probability distributions. The study focuses on nonsmooth composite…

  2. RESEARCH · CL_200031 ·

    New theory explains curvature in Probabilistic Circuits

    Researchers have developed a compositional theory for understanding curvature in Probabilistic Circuits (PCs), a type of generative model. They demonstrated that the Hessian trace, a measure of loss-surface curvature, c…

  3. RESEARCH · CL_180783 ·

    Generalized Quadratic Gradient framework unifies optimization methods

    Researchers have introduced Generalized Quadratic Gradient (GQG), a novel optimization framework that unifies and extends existing second-order optimization methods. GQG abstracts the core principles of Quadratic Gradie…

  4. TOOL · CL_167607 ·

    New research links mini-batch noise to loss landscape sharpness in SGD

    A new research paper proposes that mini-batch noise during Stochastic Gradient Descent (SGD) training influences the sharpness of the loss landscape by causing fluctuations within the dominant subspace. The authors argu…

  5. TOOL · CL_167288 ·

    New kinetic theory formalizes zeroth-order Newton methods

    Researchers have developed a formal kinetic theory for zeroth-order Newton-type methods, which are useful when gradients and Hessians are unavailable. The framework includes a Gaussian-Stein correction to accurately est…

  6. TOOL · CL_164985 ·

    Quadratic Model Shows Predictive Power for LLM Optimization Dynamics

    A new paper published on arXiv proposes that the simple quadratic model can be a surprisingly accurate predictor of optimization dynamics in large language models (LLMs). Researchers demonstrated that by analyzing the H…

  7. TOOL · CL_160971 ·

    New C-PTQ method enhances multimodal LLM quantization efficiency

    Researchers have developed C-PTQ, a novel post-training quantization method designed to improve the efficiency of multimodal large language models (MLLMs). This technique addresses performance degradation caused by outl…

  8. RESEARCH · CL_156366 ·

    Neural network research reveals functional equivalence and geometric diversity

    A new research paper explores the concept of functional equivalence in neural networks, building upon the Universal Approximation Theorem. The study reveals that multiple neural network configurations can achieve identi…

  9. TOOL · CL_141220 ·

    New Gaussian Invariant MCMC methods boost statistical efficiency

    Researchers have developed novel sampling methods, including Gaussian invariant versions of Random Walk Metropolis (RWM), Metropolis-adjusted Langevin algorithm (MALA), and a second-order Hessian or Manifold MALA. These…

  10. TOOL · CL_135386 ·

    Neural network Hessian eigenvalues explained by approximate symmetries

    Researchers have developed a new explanation for the numerous near-zero eigenvalues of the Hessian matrix in neural networks. They propose that these vanishing eigenvalues stem from approximate symmetries within the net…

  11. RESEARCH · CL_135235 ·

    Gradient Descent Theory Extended for Complex Minima and Vector Outputs

    This paper extends the theory of gradient descent (GD) with large step sizes to more complex scenarios. It addresses overparameterized least-squares problems with vector-valued outputs and analyzes neighborhoods of mani…

  12. TOOL · CL_121437 ·

    New research details Zeroth-Order optimization stability in deep learning

    A new research paper explores the stability dynamics of Zeroth-Order (ZO) optimization methods, particularly in the context of deep learning. The study identifies a specific step size condition that governs the linear s…

  13. TOOL · CL_117145 ·

    Genetic algorithms mimic clipped gradient descent in high-dimensional AI search

    Researchers have demonstrated that genetic algorithms can effectively function as a form of clipped gradient descent in high-dimensional search spaces. This process involves mutation-selection mechanisms that implicitly…

  14. RESEARCH · CL_117170 ·

    New CWGD method improves optimization noise measurement for deep learning

    Researchers have developed a new method called Curvature-Weighted Gradient Diversity (CWGD) to better measure optimization noise in deep learning models. Unlike traditional methods that treat all parameter directions eq…

  15. RESEARCH · CL_117383 ·

    Hessian Eigenvector Dynamics Reveal Optimizer Differences in Neural Network Training

    Researchers have analyzed the evolution of Hessian eigenvectors during neural network training, revealing distinct behaviors between different optimizers. The study found that SGD tends to stabilize leading curvature di…

  16. RESEARCH · CL_115597 ·

    New method offers second-order KKT guarantees for Bregman ADMM

    Researchers have developed a novel approach to analyze Bregman ADMM for nonconvex and non-Lipschitz optimization problems. This method replaces the standard Lipschitz gradient assumption with a two-sided relative smooth…

  17. RESEARCH · CL_93643 ·

    New research frameworks model gradient descent at the edge of stability

    Two new research papers explore the phenomenon of gradient descent operating at the edge of stability (EoS) in deep learning. The first paper introduces 'Edge Flow,' a system of differential equations that models gradie…

  18. RESEARCH · CL_93696 ·

    New 'Architecture Warm-Up' Stabilizes Transformer Training

    Researchers have developed a new method to stabilize the training of large Transformer models, which are often prone to instability and divergence. The approach, called "architecture warm-up," involves progressively inc…

  19. RESEARCH · CL_90893 ·

    New optimization techniques emerge for faster, more efficient AI model training · 8 sources tracked

    Several recent arXiv papers explore advancements in optimization techniques for machine learning. Researchers have proposed new methods like Weight Adaptation ASNG (WA-ASNG) to improve parallel performance in evolutiona…

  20. RESEARCH · CL_08352 ·

    New research explores how network symmetry aids optimization in overparameterized deep learning models.

    A new paper analyzes how overparameterization in neural networks aids optimization by introducing additional symmetries. These symmetries act as a form of preconditioning on the Hessian, leading to better-conditioned mi…