PulseAugur
EN
LIVE 09:10:52
ENTITY Sparse Autoencoders

Sparse Autoencoders

PulseAugur coverage of Sparse Autoencoders — every cluster mentioning Sparse Autoencoders across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
29
89 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
28
88 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
TIMELINE
  1. 2026-05-25 research_milestone Researchers published a paper detailing a new method for multilingual language steering in LLMs using sparse autoencoders. source
  2. 2026-05-21 research_milestone Researchers published a paper detailing a new method for multilingual steering in LLMs using sparse autoencoders. source
SENTIMENT · 30D

14 day(s) with sentiment data

RECENT · PAGE 1/5 · 89 TOTAL
  1. TOOL · CL_196113 ·

    New method uses Koopman operator for model interpretability

    Researchers have developed a new method for mechanistic interpretability called "Intrinsic Structure" that uses the Koopman operator to analyze the spectral properties of a model's internal dynamics. This approach aims …

  2. TOOL · CL_195936 ·

    New metric measures semantic abstractness of LLM features

    Researchers have introduced a new metric called Feature Nonlocality (FNL) to better understand the semantic abstractness of features within Sparse Autoencoders (SAEs) used in Large Language Models (LLMs). FNL measures t…

  3. TOOL · CL_195168 ·

    HyperSAE uses Poincaré geometry to boost Sparse Autoencoder performance

    A new PyTorch library called HyperSAE has been developed to improve the efficiency of Sparse Autoencoders (SAEs) by employing Poincaré hyperbolic geometry. This approach addresses the limitations of standard SAEs, which…

  4. TOOL · CL_193693 ·

    AI models' internal representations of personas analyzed

    Researchers have explored how large language models internally represent different speaking entities, such as the AI assistant, a role-playing persona, or a narrative character. By analyzing user-expressed emotions and …

  5. TOOL · CL_193672 ·

    New Sparse Autoencoders Enhance Video Representation Interpretability

    Researchers have developed spatio-temporal sparse autoencoders (SAEs) to improve the interpretability and temporal coherence of video representations. Standard SAEs, while good at decomposing features, often sacrifice t…

  6. TOOL · CL_193597 ·

    New MMDiff framework enhances control and interpretability of multimodal LLMs

    Researchers have developed MMDiff, a novel framework designed to enhance the interpretability and control of Multimodal Large Language Models (MLLMs). This system trains multimodal sparse autoencoders (SAEs) to identify…

  7. RESEARCH · CL_191390 ·

    New research probes Sparse Autoencoders for neural network interpretability

    Two new research papers explore the interpretability of neural networks, specifically focusing on Sparse Autoencoders (SAEs). The first paper questions the effectiveness of SAEs in capturing human-like category boundari…

  8. TOOL · CL_187401 ·

    New CircuitSteer framework enhances LLM control via multi-layer semantic circuits

    Researchers have developed a new framework called CircuitSteer, which uses Sparse Autoencoders to identify and manipulate specific semantic circuits within multiple layers of large language models. This method allows fo…

  9. TOOL · CL_187254 ·

    New method probes bias in AI L2 speaking assessment systems

    Researchers have developed a new method to analyze bias in AI systems used for second language (L2) speaking assessments. This approach utilizes Concept Activation Vectors (CAVs) to probe how models like BERT and Whispe…

  10. TOOL · CL_178048 ·

    New GLASS method injects AI text style without retraining

    Researchers have introduced GLASS, a novel method detailed in an arXiv preprint that allows AI text generation to adopt a specific writing style without requiring model retraining or retrieval of external data. This tec…

  11. TOOL · CL_174313 ·

    New framework explains image similarity using concept activation vectors

    Researchers have developed a new framework to explain image similarity using automatically extracted Concept Activation Vectors (CAVs). This model-agnostic approach utilizes Sparse Autoencoders (SAEs) to identify concep…

  12. TOOL · CL_171971 ·

    Interpretable AI agents in GOHR game revealed by sparse autoencoders

    Researchers have developed interpretable experiments for a Transformer agent operating in the Game of Hidden Rules (GOHR). By training sparse autoencoders on the agent's decision-token embeddings, they were able to reco…

  13. RESEARCH · CL_167444 ·

    New framework analyzes sparse autoencoder feature effects

    Researchers have introduced Feature-Effect Geometry Analysis (FEGA), a new framework for understanding the downstream effects of sparse autoencoder (SAE) features. FEGA analyzes how interventions on SAE features alter m…

  14. TOOL · CL_165006 ·

    New training method enhances LLM interpretability by reducing signal loss

    Researchers have developed a new method called replacement-aware training to improve the interpretability of large language models. This technique trains sparse auto-encoders (SAEs) to be robust to errors introduced by …

  15. TOOL · CL_160880 ·

    Microsoft's Aurora AI model shows statistical learning, not mechanistic understanding of chemistry

    Researchers have investigated Microsoft's Aurora model, a foundation model fine-tuned for atmospheric chemistry, to understand its internal learning mechanisms. While the model demonstrates skill in predicting air quali…

  16. TOOL · CL_160737 ·

    New ParityTransformer architecture enables scalable, interpretable AI models

    Researchers have developed the ParityTransformer, a novel GPT-2 scale architecture designed for enhanced interpretability. This model utilizes a Deep Parity Bottleneck (DPB) at each layer, which replaces expensive learn…

  17. TOOL · CL_156417 ·

    New Hyperdimensional Probe Decodes LLM Representations

    Researchers have developed a new interpretability method called the Hyperdimensional Probe, which combines symbolic representations with neural probing to better understand the internal workings of Large Language Models…

  18. RESEARCH · CL_154406 ·

    Sparse Autoencoders Offer Interpretable Insights into LLM Data and Behavior · 4 sources tracked

    Researchers are exploring the use of sparse autoencoders (SAEs) as a more cost-effective and interpretable method for analyzing large-scale text corpora and understanding the internal workings of large language models. …

  19. TOOL · CL_154396 ·

    New FAST training method enhances Sparse Autoencoders for instruct models

    Researchers have developed a new training paradigm called Finetuning-aligned Sequential Training (FAST) to improve Sparse Autoencoders (SAEs) for instruct models. Traditional block training methods introduce gradient no…

  20. TOOL · CL_154248 ·

    New metric measures monosemanticity in AI explanations

    Researchers have developed a new metric called the Tversky Monosemanticity Score (TMS) to better assess the quality of explanations generated by Sparse Autoencoders (SAEs) in mechanistic interpretability. Unlike previou…