A new open-source model called DWARF-55M-Base has been released, featuring a novel sparse attention architecture. This architecture, named Dynamic Sparse Query-Gather (DSQG), aims to reduce computational costs by sampling tokens rather than attending to every previous token. The model is trained on 10 billion tokens and is available under an Apache 2.0 license, with experimental code for further exploration. AI
IMPACT Introduces a new sparse attention mechanism that could lead to more efficient LLM architectures.
RANK_REASON Release of a new open-source model with a novel architecture. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →