Lindsey et al.
PulseAugur coverage of Lindsey et al. — every cluster mentioning Lindsey et al. across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New training method enhances LLM interpretability by reducing signal loss
Researchers have developed a new method called replacement-aware training to improve the interpretability of large language models. This technique trains sparse auto-encoders (SAEs) to be robust to errors introduced by …
-
New 'prolepsis' phenomenon identified in small transformer models
Researchers have identified a phenomenon called 'prolepsis' in small transformer models, where the model commits to a decision early in its processing and cannot correct it. This commitment is sustained by task-specific…
-
Language Model Neurons Found to Be Sparse, Aiding Interpretability
Researchers have demonstrated that the neurons within a language model's MLP layers exhibit a degree of sparsity comparable to that of Sparse Autoencoders (SAEs). This finding enables the development of a gradient-based…