Researchers have introduced Sparse Delta Memory (SDM), a novel architecture designed to enhance the long-context recall capabilities of linear RNNs. By employing a sparse addressing scheme, SDM significantly increases the hidden state capacity of gated linear RNNs, outperforming traditional transformer architectures on in-context learning and long-context retrieval tasks under similar computational constraints. The architecture extends the Gated DeltaNet by utilizing sparse reads and writes to an explicit memory, and can further improve performance on common-knowledge and reasoning tasks when its initial state is learned as a parametric memory. AI
IMPACT Enhances long-context recall in RNNs, potentially offering an alternative to transformer architectures for specific tasks.
RANK_REASON The cluster describes a new research paper detailing a novel architecture for RNNs.
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gated DeltaNet
- Gotit.pub
- Hugging Face
- IArxiv
- Linear RNNs
- ScienceCast
- Sparse Delta Memory
- transformer architectures
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →