English(EN)Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
新研究探索稀疏自编码器在人工智能可解释性和泛化方面的应用
作者PulseAugur 编辑部·[7 个来源]·
研究人员正在探索稀疏自编码器(SAEs)来解释复杂的语言和视觉模型。一篇论文介绍了用于各种Qwen3模型尺寸的Qwen3-Instruct SAEs,展示了它们在引导模型行为方面的应用。另一项研究调查了SAEs如何揭示Transformer泛化的局限性并提高对分布外输入的鲁棒性。第三篇论文提出新的稀疏正则化器来增强Top-k SAEs的可解释性,表明它们可以补充架构稀疏性。最后,提出了一个使用概念标注和合成基准来评估SAE可解释性的框架,表明适中的字典尺寸可以产生最可解释的SAEs。
AI
arXiv:2606.26620v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features. However, training SAEs is computationally expensive, and available open-s…
arXiv:2606.26396v1 Announce Type: new Abstract: Pre-trained transformers have demonstrated remarkable generalization abilities, at times extending beyond the scope of their training data. Yet, real-world deployments often face unexpected or adversarial data that diverges from tra…
arXiv cs.AI
TIER_1English(EN)·Nathana\"el Jacquier, Maria Vakalopoulou, Mahdi S. Hosseini·
arXiv:2606.27321v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have become a leading tool for interpreting the representations of vision foundation models, decomposing their polysemantic activations into a larger set of sparse, more monosemantic features. The Top-$k…
Sparse autoencoders (SAEs) have become a leading tool for interpreting the representations of vision foundation models, decomposing their polysemantic activations into a larger set of sparse, more monosemantic features. The Top-$k$ SAE, a now-standard variant, enforces sparsity a…
Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features. However, training SAEs is computationally expensive, and available open-source SAE models remain limited. In this work, we …
arXiv cs.AI
TIER_1English(EN)·Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Beg\"um Demir·
arXiv:2606.24716v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuri…
Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence. We present a human-gro…