A new technical report introduces Instella-MoE, an open-source Mixture-of-Experts (MoE) language model with 16 billion total parameters. Trained on AMD Instinct GPUs, the model incorporates innovations like Gated Multi-head Latent Attention and FarSkip-Collective connectivity for efficient training and inference. Instella-MoE demonstrates strong performance on pre-training benchmarks, outperforming other open models, and achieves competitive results on instruction-following, reasoning, math, coding, and chat tasks. The developers are releasing the model weights, training configurations, data mixtures, and code to foster reproducible research. AI
IMPACT Provides a new, fully open foundation for efficient MoE models and reproducible research in the field.
RANK_REASON The cluster contains a technical report detailing a new open-source language model release. [lever_c_demoted from research: ic=1 ai=1.0]
- AMD Instinct MI300X
- AMD Instinct MI325X
- Instella-MoE
- Moonlight-16B-A3B
- OLMo-3-7B
- OLMoE-1B-7B
- Qwen3.5-4B
- SmolLM3-3B
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →