Three new arXiv papers explore the geometric and statistical properties of token embeddings in language models. The first paper identifies a "hub of short rows" near the origin of token embedding tables that inflates intrinsic dimension estimates, showing that removing this hub leads to more consistent dimension readings across models like GPT-2, K3, and GLM-4.7. The second paper introduces the "Context Staircase" concept, describing how embeddings progressively learn more complex, context-dependent statistical signatures as training advances. The third paper develops a statistical framework linking token prediction to representation geometry, demonstrating how prediction accuracy and recovered geometry translate into downstream task performance. AI
IMPACT These papers offer a deeper understanding of how language models learn representations, potentially guiding future architectural and training improvements.
RANK_REASON The cluster consists of three academic papers published on arXiv detailing theoretical findings about token embeddings in language models.
- arXiv
- Context Staircase
- GLM-4.7
- GPT-2
- Hellinger distance
- Hugging Face
- K3
- language models
- Pythia
- token embeddings
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →