Researchers have introduced a new optimization scheme for deep neural networks that utilizes a dynamic $\ell_p$-norm, moving beyond the limitations of fixed $\ell_2$ and $\ell_\infty$ norms. This novel approach, termed LPSGD and LPSGDM, aims to improve convergence and generalization by adapting the norm's parameter $p$ throughout the training process. The method begins with a large $p$ to manage high-curvature directions and gradually decreases $p$ towards 2 for more stable updates, theoretically achieving an $O(T^{-1/2})$ convergence rate for non-convex problems. AI
IMPACT Introduces a novel optimization technique that could improve training efficiency and generalization for deep learning models.
RANK_REASON The cluster contains a research paper detailing a new theoretical scheme and experimental results for deep neural network optimization.
- CIFAR-10
- CIFAR-100
- Deep Neural Networks
- ImageNet-1K
- ResNet-18
- ResNet-50
- SGD
- SGD with momentum
- VGG-11
- LPSGD
- LPSGDM
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →