Pythia 1B
PulseAugur coverage of Pythia 1B — every cluster mentioning Pythia 1B across labs, papers, and developer communities, ranked by signal.
-
New RL method slashes LLM pretraining time by 66%
Researchers have developed AC-ODM, a novel method that uses reinforcement learning to optimize the composition of pretraining data for large language models (LLMs). This approach significantly improves sample efficiency…
-
New method validates LLM circuits using ablation tests
Researchers have developed a new method for discovering circuits within large language models by clustering attention head co-activation statistics. This approach, termed "closure-validated circuit discovery," uses caus…
-
Researchers track attention circuit formation in 1B-class language models
A new research paper investigates the emergence of attention circuits in language models, specifically tracking how different types of attention heads form across various model architectures and training datasets. The s…
-
New Sparse Autoencoder Model Enhances LLM Feature Interpretability
Researchers have introduced the Sign-Aware Gated Sparse Autoencoder (SA-GSAE), a novel architecture designed to improve the interpretability of features extracted from Large Language Models. Unlike standard SAEs that en…