Gao et al. (2024)
PulseAugur coverage of Gao et al. (2024) — every cluster mentioning Gao et al. (2024) across labs, papers, and developer communities, ranked by signal.
-
New Study Finds Up to 77% of AI Features May Be Non-Functional
A new study has revealed that a significant portion of features identified by Sparse Autoencoders (SAEs), a tool used in mechanistic interpretability, may not actually be functional. The research found that up to 77% of…
-
New research questions Sparse Autoencoder interpretability and introduces new evaluation benchmark
Two new research papers investigate the effectiveness and interpretability of Sparse Autoencoders (SAEs), a standard method for decomposing neural representations. The first paper, "From Geometric Recovery to Causal Val…
-
EleutherAI releases open-source tool for interpreting AI model features
EleutherAI has released an open-source library for automatically interpreting features within sparse autoencoders, a method used to decompose model activations. This tool leverages large language models like Llama 3.1 a…