Researchers have demonstrated that the geometric structures observed in language model embeddings, such as circles for months and saddle-shaped manifolds for chords, are directly predictable from the statistical symmetries present in the training data. By applying harmonic analysis and group theory, the study shows that if a word family's co-occurrence statistics are invariant under a specific group G, the learned embeddings will correspond to the irreducible representations of G. This framework successfully explains and reproduces the circular geometry of months, the unified geometry of musical chords, and the spherical representation of celestial objects found in large language models. AI
IMPACT Provides a theoretical framework for understanding and predicting the geometric structure of LLM embeddings based on data symmetries.
RANK_REASON Academic paper published on arXiv detailing theoretical findings about LLM embeddings. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- Bach
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- LLMs
- Pringle
- ScienceCast
- T/Idi Primary School
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →