Researchers have developed a theoretical framework to explain the phenomenon of similar representations for translated sentences in the inner layers of multilingual language models. This framework is based on the idea that languages share abstract hierarchical structures while surface-level details are language-specific. By generating synthetic languages with shared upper-level but distinct lower-level structures, the study found that belief propagation, when encoded in successive layers, accurately predicts the behavior of transformers trained on similar data. The theory accounts for cross-lingual similarity peaking in middle layers, its coexistence with language-specific structures, and its enhancement with factors like language proximity and model quality. AI
IMPACT Provides a theoretical explanation for cross-lingual similarities in LLMs, potentially guiding future model development.
RANK_REASON The item is a research paper published on arXiv detailing a new theoretical framework for understanding language models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- A theory of platonic representations in language models
- Bayes-optimal next-token predictor
- belief propagation
- CatalyzeX Code Finder for Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- IArxiv Recommender
- Platonic Representation Hypothesis
- ScienceCast
- transformers
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →