A new research paper investigates the cross-lingual validity of Sparse Autoencoder (SAE) features in Google's Gemma 2 and Gemma 3 language models. The study found that while many features appear frequently across different languages, they often have minimal or inconsistent causal effects on translation tasks. However, one specific feature in both Gemma 2 and Gemma 3 consistently improved translation quality as measured by COMET scores when amplified, and degraded it when ablated, suggesting a language-agnostic translation direction. AI
IMPACT Highlights limitations in current methods for interpreting LLM features across languages, suggesting a need for more robust validation techniques.
RANK_REASON Research paper published on arXiv detailing findings about language model features. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →