Researchers have introduced WhiteMatter, a novel architecture for Transformers that enhances cross-layer connections. Unlike traditional Transformers where each layer only accesses KV pairs from its own depth, WhiteMatter allows every attention layer to connect with representations from all layers of past tokens. This is achieved through a router that mixes layer states into KV channels, enabling consumer layers to attend to these channels. The number of channels, k, controls the KV-cache size, offering potential memory footprint reduction. Experiments show WhiteMatter outperforms a standard Transformer with more layers and maintains significant gains even with KV-cache compression. AI
IMPACT Introduces a novel Transformer architecture that could improve efficiency and performance in large language models.
RANK_REASON The cluster contains a research paper detailing a new Transformer architecture. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- ScienceCast
- Transformer
- WhiteMatter
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →