A new research paper proposes that the weight magnitudes in trained transformers can be described by a Weibull distribution. The study identifies a pre-training statistic, the bigram conditional entropy, as a key predictor for the growth of the scale parameter in this distribution. This predictive law holds across various learning rates and model architectures, suggesting a fundamental relationship between data predictability and model weight scaling during training. AI
IMPACT This research offers a new theoretical framework for understanding and potentially predicting transformer training dynamics based on data properties.
RANK_REASON The cluster contains a single arXiv paper detailing a new research finding about transformer training. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →