Sparse Autoencoders
PulseAugur coverage of Sparse Autoencoders — every cluster mentioning Sparse Autoencoders across labs, papers, and developer communities, ranked by signal.
- instance of Saessolsheim 90%
- used by Vision Transformers 90%
- used by Gotit.pub 80%
- developed by Gotit.pub 70%
- used by Saessolsheim 70%
- used by UniPro 70%
- used by GPT-2 small 70%
- used by InternVL3.5 70%
- used by CatalyzeX Code Finder for Papers 70%
- instance of Linear Representation Hypothesis 60%
- affiliated with GPT-2 small 50%
- affiliated with Vision Transformers 50%
- 2026-05-25 research_milestone Researchers published a paper detailing a new method for multilingual language steering in LLMs using sparse autoencoders. source
- 2026-05-21 research_milestone Researchers published a paper detailing a new method for multilingual steering in LLMs using sparse autoencoders. source
8 day(s) with sentiment data
-
New methods improve harmful meme detection in vision-language models
Researchers have developed new methods to improve the detection of harmful memes by vision-language models. One approach, "Decodable but Misrouted," uses sparse autoencoders and causal interventions to identify whether …
-
New framework PhysSAE enhances interpretability of physics-informed neural networks
A new framework called PhysSAE has been developed for mechanistic interpretability of Physics-Informed Neural Networks (PINNs). This framework uses overcomplete sparse autoencoders to analyze the internal representation…
-
MonoTM framework enhances topic modeling with interpretable features
Researchers have introduced MonoTM, a novel framework designed to enhance topic modeling by extracting interpretable monosemantic features. This approach separates the estimation of document-topic mixtures from the sema…
-
New methods enhance LLM unlearning for improved safety and utility
Researchers have developed new methods for machine unlearning in Large Language Models (LLMs) to remove unwanted knowledge while preserving model utility. One approach, Dynamic Sparse Autoencoder Guardrails (DSG), uses …
-
New theory quantifies neural network feature superposition limits
Researchers have developed a new theoretical framework to understand feature superposition in neural networks, addressing the issue of cross-feature interference. By modeling linear accessibility as a compressed sensing…
-
Research probes stereotype representation in multilingual LLMs
A new research paper investigates how stereotypes manifest within multilingual large language models (LLMs). The study compares various methods like linear probing and sparse autoencoders across models such as Llama-3.1…
-
Gemma-2-9B-IT model shows language-agnostic informal register
Researchers have identified and controlled a language-agnostic informal register within the Gemma-2-9B-IT model. Using Sparse Autoencoders (SAEs) across English, Hebrew, and Russian, they discovered a robust cross-lingu…
-
New method uses sparse autoencoders for causal adjustment in text data
Researchers have proposed a new method for adjusting causal questions in text data using sparse autoencoders (SAEs). This approach aims to balance the need for dense representations to capture confounding variables with…
-
New KronSAE design enhances sparse autoencoder efficiency and interpretability
Researchers have introduced KronSAE, a novel design for Sparse Autoencoders (SAEs) that improves their efficiency and interpretability. Unlike traditional SAEs that treat latent dictionaries as flat coordinates, KronSAE…
-
New NIFS method enhances LLM steering via sparse autoencoders
Researchers have developed a new method called Neighbor Integrated Feature Selection (NIFS) to improve the effectiveness of steering large language models using sparse autoencoders (SAEs). Traditional methods select fea…
-
New EXPOSE framework enhances VFM explainability in pathology
Researchers have introduced EXPOSE, a novel framework designed to enhance the explainability and domain robustness of Vision Foundation Models (VFMs) in computational pathology. By employing Sparse Autoencoders (SAEs), …
-
New method uses geometry to interpret protein language model features
Researchers have developed a new method to interpret the latent features within protein language models (pLMs) by using geometric annotations of protein backbones. This approach, applied to the ESM-2 model, reveals that…
-
Sparse autoencoders show reward filtering captures solution completeness, not reasoning quality
Researchers have developed a reward-informed sparse autoencoder (RI-SAE) to interpret language model activations, specifically focusing on reasoning capabilities. While the RI-SAE successfully separated high-reward and …
-
New CAST framework enhances audibility of clinical AI models
Researchers have developed a new framework called CAST (Concept-guided Artifact Suppression Tuning) to make clinical language models more auditable and robust. This method uses Sparse Autoencoders to identify and suppre…
-
New method disentangles AI model features for improved multi-task merging
Researchers have developed a novel framework for merging multiple AI models into a single, more capable generalist model. This method addresses the challenge of "superposition," where task-specific features become entan…
-
New method visualizes feature evolution in Vision Transformers
Researchers have developed a new method to visualize the training process of Vision Transformers (ViTs) by examining feature evolution across network depth and training time. Using Sparse Autoencoders (SAEs), they can i…
-
New method enhances AI interpretability for extreme Earth event prediction
Researchers have developed a new method called SAE-Xplainers to improve the interpretability of deep learning models used for analyzing extreme Earth events. This approach modifies Sparse Autoencoders (SAEs) to better c…
-
SpIn-ViT framework enhances Vision Transformer interpretability and accuracy
Researchers have developed SpIn-ViT, a novel framework that integrates a Vision Transformer (ViT) with a Sparse Autoencoder (SAE) for enhanced interpretability and performance. Unlike previous methods that applied SAEs …
-
Research paper flags critical flaw in language model interpretation methods
A new research paper highlights a critical flaw in the evaluation of sparse autoencoders (SAEs), a method used to interpret language models. The study reveals that the standard practice of measuring a latent's importanc…
-
New Framework Generates Natural-Language Explanations for LLM Features
Researchers have developed SAEVerbalizer, a new framework designed to generate natural-language explanations for features extracted by Sparse Autoencoders (SAEs) from large language models (LLMs). Current methods for ex…