PulseAugur
EN
LIVE 00:47:07
ENTITY Hessian

Hessian

PulseAugur coverage of Hessian — every cluster mentioning Hessian across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
3
18 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
3
18 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

3 day(s) with sentiment data

RECENT · PAGE 1/2 · 26 TOTAL
  1. TOOL · CL_268915 ·

    New research details quantization failures in looped transformers

    Researchers have identified two critical failure modes in post-training quantization (PTQ) for looped transformers, which reuse weights across recurrence steps. The first, termed 'feedback exposure,' occurs when a quant…

  2. TOOL · CL_257077 ·

    New algorithm simplifies federated bilevel optimization using first-order gradients

    Researchers have developed a new algorithm for federated stochastic bilevel optimization that avoids the need for computationally expensive second-order Hessian and Jacobian matrices. This novel approach, termed federat…

  3. RESEARCH · CL_257091 ·

    New research explores geometric optimization at the 'Edge of Stability' in associative memories

    Researchers have explored the geometric properties of high-capacity kernel logistic regression (KLR) associative memories, identifying a critical hyperparameter regime known as the "Ridge of Optimization." This region i…

  4. TOOL · CL_216001 ·

    New Newton's Method Achieves O(1/k^3) Convergence Rate

    Researchers have developed a novel direct accelerated Newton method for minimizing convex functions with Lipschitz continuous Hessians. This new algorithm operates solely with primal variables and requires only one line…

  5. TOOL · CL_212162 ·

    New 'Rényi Sharpness' metric shows strong generalization correlation

    Researchers have introduced "Rényi sharpness," a new metric for evaluating neural network generalization that aims to improve upon existing methods. Unlike traditional sharpness measures that focus on average loss or ma…

  6. TOOL · CL_210528 ·

    New theory reframes optimizer stability in deep learning

    Researchers have identified a phenomenon in deep learning where gradient-based optimizers maintain stable Hessian eigenvalues above theoretically predicted instability thresholds. This deviation, observed to be as high …

  7. TOOL · CL_200182 ·

    New theoretical bounds improve Langevin sampling for complex distributions

    Researchers have developed new theoretical bounds for the Moreau--Yosida unadjusted Langevin algorithm (MYULA), a method used for sampling from complex probability distributions. The study focuses on nonsmooth composite…

  8. RESEARCH · CL_200031 ·

    New theory explains curvature in Probabilistic Circuits

    Researchers have developed a compositional theory for understanding curvature in Probabilistic Circuits (PCs), a type of generative model. They demonstrated that the Hessian trace, a measure of loss-surface curvature, c…

  9. RESEARCH · CL_180783 ·

    Generalized Quadratic Gradient framework unifies optimization methods

    Researchers have introduced Generalized Quadratic Gradient (GQG), a novel optimization framework that unifies and extends existing second-order optimization methods. GQG abstracts the core principles of Quadratic Gradie…

  10. TOOL · CL_167607 ·

    New research links mini-batch noise to loss landscape sharpness in SGD

    A new research paper proposes that mini-batch noise during Stochastic Gradient Descent (SGD) training influences the sharpness of the loss landscape by causing fluctuations within the dominant subspace. The authors argu…

  11. TOOL · CL_167288 ·

    New kinetic theory formalizes zeroth-order Newton methods

    Researchers have developed a formal kinetic theory for zeroth-order Newton-type methods, which are useful when gradients and Hessians are unavailable. The framework includes a Gaussian-Stein correction to accurately est…

  12. TOOL · CL_164985 ·

    Quadratic Model Shows Predictive Power for LLM Optimization Dynamics

    A new paper published on arXiv proposes that the simple quadratic model can be a surprisingly accurate predictor of optimization dynamics in large language models (LLMs). Researchers demonstrated that by analyzing the H…

  13. TOOL · CL_160971 ·

    New C-PTQ method enhances multimodal LLM quantization efficiency

    Researchers have developed C-PTQ, a novel post-training quantization method designed to improve the efficiency of multimodal large language models (MLLMs). This technique addresses performance degradation caused by outl…

  14. RESEARCH · CL_156366 ·

    Neural network research reveals functional equivalence and geometric diversity

    A new research paper explores the concept of functional equivalence in neural networks, building upon the Universal Approximation Theorem. The study reveals that multiple neural network configurations can achieve identi…

  15. TOOL · CL_141220 ·

    New Gaussian Invariant MCMC methods boost statistical efficiency

    Researchers have developed novel sampling methods, including Gaussian invariant versions of Random Walk Metropolis (RWM), Metropolis-adjusted Langevin algorithm (MALA), and a second-order Hessian or Manifold MALA. These…

  16. TOOL · CL_135386 ·

    Neural network Hessian eigenvalues explained by approximate symmetries

    Researchers have developed a new explanation for the numerous near-zero eigenvalues of the Hessian matrix in neural networks. They propose that these vanishing eigenvalues stem from approximate symmetries within the net…

  17. RESEARCH · CL_135235 ·

    Gradient Descent Theory Extended for Complex Minima and Vector Outputs

    This paper extends the theory of gradient descent (GD) with large step sizes to more complex scenarios. It addresses overparameterized least-squares problems with vector-valued outputs and analyzes neighborhoods of mani…

  18. TOOL · CL_121437 ·

    New research details Zeroth-Order optimization stability in deep learning

    A new research paper explores the stability dynamics of Zeroth-Order (ZO) optimization methods, particularly in the context of deep learning. The study identifies a specific step size condition that governs the linear s…

  19. TOOL · CL_117145 ·

    Genetic algorithms mimic clipped gradient descent in high-dimensional AI search

    Researchers have demonstrated that genetic algorithms can effectively function as a form of clipped gradient descent in high-dimensional search spaces. This process involves mutation-selection mechanisms that implicitly…

  20. RESEARCH · CL_117170 ·

    New CWGD method improves optimization noise measurement for deep learning

    Researchers have developed a new method called Curvature-Weighted Gradient Diversity (CWGD) to better measure optimization noise in deep learning models. Unlike traditional methods that treat all parameter directions eq…