This paper delves into the stability and generalization of straight-through estimators (STE) for training two-layer quantized neural networks, analyzed through the lens of Statistical Learning Theory. The research establishes that in a saturated-output regime, the STE recursion is equivalent to stochastic subgradient descent on a convex latent loss. This equivalence allows for a stability analysis, yielding explicit L2 on-average model-stability and generalization bounds. The findings provide an excess induced-risk guarantee and an optimal-order expected excess misclassification error rate under margin separability. AI
IMPACT Provides theoretical insights into the training dynamics of quantized neural networks, potentially informing future model architectures and training techniques.
RANK_REASON Academic paper on theoretical aspects of neural network training. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Statistical Learning Theory
- Stochastic Subgradient Descent
- Straight-Through Estimators
- Two-Layer Quantized Neural Networks
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →