Pythia 70M
PulseAugur coverage of Pythia 70M — every cluster mentioning Pythia 70M across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New theory explains heavy-tail emergence in neural optimizer dynamics
Researchers have developed a new method to understand how heavy-tailed spectral densities emerge in neural network weight matrices, which are indicators of implicit self-regularization. They formulated this emergence as…
-
New method measures predictability in neural network training dynamics
Researchers have developed a new method to measure structured predictability in neural network training dynamics. This approach uses complementary probe families to analyze temporal redundancy, identifying when and unde…
-
Expander SAEs offer parameter-efficient dictionaries for neural network interpretability
Researchers have introduced Expander Sparse Autoencoders (SAEs), a novel approach to interpret neural network activations by using parameter-efficient dictionaries. This method significantly reduces the number of learne…
-
Weibull framework reveals AdamW training dynamics in transformers
A new research paper explores the evolution of weight-scale parameters in transformer models during AdamW training. The study derives a three-force decomposition of the squared weight norm, identifying alignment, inject…
-
Researchers find independently trained transformers compute same function via random rotation
Researchers have discovered a phenomenon called "polymorphism" in independently trained transformers, where they compute the same function but use different internal coordinate systems that are rotated versions of each …
-
New methods enhance sparse autoencoder interpretability and stability
Researchers have developed new methods to address limitations in sparse autoencoders (SAEs), which are used to interpret the internal representations of large language models. One paper introduces adaptive elastic net S…