A new research paper proposes an alternative to the standard geometric derivation of neural scaling exponents, which typically relies on the intrinsic dimension of a data manifold. The authors demonstrate that for modular addition in $\mathbb{Z}_p$, this standard method is undefined because the exact algebraic solution involves an orbit of $\mathbb{Z}_p$ acting by isometries. Instead, they show that the relationship follows an exponential curve related to the hidden width of the network, with a high R-squared value, suggesting this model better captures the observed behavior. AI
IMPACT Proposes a new theoretical framework for understanding neural scaling, potentially impacting model design and analysis.
RANK_REASON The cluster contains a single academic paper detailing a novel theoretical approach to neural scaling exponents. [lever_c_demoted from research: ic=1 ai=1.0]
- data manifold
- Frédéric Cadet
- hidden width
- intrinsic dimension
- \mathbb{Z}_p
- modular addition
- neural scaling exponents
- Tikhonov regularization
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →