A new research paper published on arXiv explores the distinction between decodability and causality in language models. The study introduces a method to decompose probe readouts into sparse autoencoder (SAE) features, ranking them by alignment with probe data and gradient sensitivity to model behavior. This decomposition reveals that features aligned with probe geometry do not necessarily causally drive model behavior, with interventions showing significant differences in behavioral impact. AI
IMPACT Introduces a new methodology for analyzing language model behavior, potentially improving the understanding of model decision-making processes.
RANK_REASON The cluster contains a single research paper published on arXiv detailing a new methodology for analyzing language model behavior. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- Buerger et al.
- Camille Davis
- CatalyzeX
- DagsHub
- Gemma2-9B-Instruct
- Gotit.pub
- Hugging Face
- Long et al.
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →