Two new research papers explore the internal workings of multilingual large language models (mLLMs) and their translation capabilities. The first paper questions the concept of a single "lingua franca" within these models, finding that different probing methods yield conflicting results about language-specific states. The second paper proposes a more modular view of translation, suggesting that models first establish target word order before generating the surface form of the target language, with specific attention heads dedicated to syntactic transformations. AI
IMPACT These studies offer deeper insights into how multilingual models process and translate languages, potentially guiding future model development and evaluation.
RANK_REASON Two academic papers published on arXiv detailing new research into the internal mechanisms of multilingual LLMs.
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- English
- Gaussian mixture model
- Gotit.pub
- Hugging Face
- Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs
- Litmaps
- Multilingual Large Language Models
- ScienceCast
- scite Smart Citations
- Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →