A new paper from arXiv questions the validity of using cosine similarity to measure the transferability of interpretability artifacts under quantization. The authors argue that reported statistics lack the necessary noise floor information for proper interpretation. They propose a method to measure this noise floor, demonstrating on Qwen2.5-1.5B-Instruct that at INT4 quantization, the direction of interpretability artifacts actually rotates, a finding obscured by standard reporting practices. The paper also highlights that scale-invariant statistics cannot differentiate between translation and attenuation of a transferred decision variable, suggesting alternative reporting recommendations. AI
IMPACT Challenges current methods for evaluating AI interpretability, potentially impacting how models are assessed for safety and robustness.
RANK_REASON Academic paper published on arXiv detailing novel research findings. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →