PulseAugur
EN
LIVE 07:58:37

Neural network generalization near interpolation analyzed via statistical mechanics · 2 sources tracked

Two new arXiv papers explore the behavior of shallow neural networks with extensive width, focusing on their generalization capabilities near the interpolation threshold. The research analyzes these networks using statistical mechanics, revealing a phase transition between a universal phase where generalization error is independent of weight distribution and a specialization phase where it becomes dependent. The findings suggest that while highly predictive solutions exist near interpolation, practical algorithms may struggle to find them due to statistical-to-computational gaps. AI

IMPACT Provides theoretical insights into neural network behavior, potentially informing future model architectures and training strategies.

RANK_REASON Two academic papers published on arXiv detailing theoretical analysis of neural network generalization.

Read on arXiv stat.ML →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Neural network generalization near interpolation analyzed via statistical mechanics · 2 sources tracked

COVERAGE [2]

  1. arXiv stat.ML TIER_1 English(EN) · Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk ·

    Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation

    arXiv:2501.18530v3 Announce Type: replace Abstract: We consider a teacher-student model of supervised learning with a fully-trained two-layer neural network whose width $k$ and input dimension $d$ are large and proportional. We provide an effective theory for approximating the Ba…

  2. arXiv stat.ML TIER_1 English(EN) · Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk ·

    Statistical mechanics of extensive-width Bayesian neural networks near interpolation

    arXiv:2505.24849v2 Announce Type: replace Abstract: For three decades statistical mechanics has been providing a framework to analyse neural networks. However, the theoretically tractable models, e.g., perceptrons, random features models and kernel machines, or multi-index models…