Researchers have developed a theoretical framework to understand the relationship between activation patching and weight-space ablation, two methods used to determine causal responsibility in neural networks. The theory, tested on idealized models and small transformers, reveals conditions under which these methods agree and when they diverge, particularly concerning how they measure changes in model outputs. The findings suggest that while patching measures the contrast in activations, ablation measures absolute levels, leading to differing conclusions about component importance. AI
IMPACT Provides a theoretical basis for interpreting causal attributions in neural networks, potentially improving model interpretability.
RANK_REASON Academic paper detailing a new theoretical framework for analyzing neural network components. [lever_c_demoted from research: ic=1 ai=1.0]
- activation patching
- arXiv
- A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation
- Hugging Face
- multilayer perceptron
- Spearman
- weight-space ablation
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →