A new research artifact, Aether-6B-11Attn-base, has been released mid-training, offering a unique look into the development process of a model that combines eleven heterogeneous mixers including conventional attention, state-space, convolutional, and linear-time components. This model, characterized by its deep and narrow architecture of 121 layers with 5.79B parameters, is not optimized for serving and exhibits slow token generation. The release emphasizes transparency, acknowledging that the training code was not archived and the model was reconstructed from a sharded checkpoint, with validation performed through parameter matching and perplexity checks. AI
IMPACT Provides a rare look into the mid-training state of a novel heterogeneous model architecture, offering insights for researchers studying model development and component integration.
RANK_REASON The item describes the release of a research artifact (a mid-training model checkpoint) with a focus on its architectural experimentation and transparent disclosure of its development process, fitting the research category. [lever_c_demoted from research: ic=1 ai=1.0]
- Aether-6B-11Attn-base
- Aether-7B-5Attn
- Apache Software License 2.0
- ETNews
- FSDPC
- Mamba
- Transformer++
- VIDRAFT
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →