A new research paper published on arXiv investigates the reliability of attention-head ablations in language models for making causal claims about component functions. The study, using GPT-2 small and DistilGPT2, demonstrates that common implementation methods for 'zeroing' attention heads can yield results that are nearly uncorrelated with corrected pre-projection ablations. The research highlights that evaluation metrics like binary accuracy can obscure effects at behavioral extremes, while gold-token log-probability offers a more graded measure. By employing matched controls and a discovery/held-out split, the paper shows that corrected per-head effect rankings are highly stable, but evidence for task specificity remains weak. AI
IMPACT Highlights potential flaws in common methods for understanding internal language model workings, suggesting a need for more rigorous analysis.
RANK_REASON Research paper published on arXiv detailing methodology for analyzing language model components. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →