A new research paper explores how Stochastic Gradient Descent (SGD) selects specific solutions when learning the identity function in deep linear residual networks. While many solutions exist that minimize population loss, SGD consistently favors particular ones, which can be understood through the lens of entropic loss. This entropic term, which penalizes the expected squared norm of the minibatch gradient, helps distinguish between different functional decompositions of the identity across network layers. The study analytically characterizes the minimizers of this entropic loss and uses these predictions to explain the observed behavior of SGD-trained networks. AI
IMPACT Provides theoretical insights into the optimization dynamics of deep learning models.
RANK_REASON Academic paper on a specific machine learning optimization technique. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →