AMD has released Instella-MoE-16B-A3B, an open-source Mixture-of-Experts language model. This model features 16 billion total parameters but only activates 2.8 billion per token, utilizing architectural innovations like Gated Multi-head Latent Attention and FarSkip-Collective connectivity for improved training and inference speeds. While the training code is MIT licensed and highly reusable, the model weights are released under a ResearchRAIL license, restricting commercial use. Instella-MoE-16B-A3B demonstrates strong performance, leading fully open models on several benchmarks and achieving a 64K context window. AI
IMPACT Provides a new open-source MoE model for research, potentially advancing understanding of expert routing and long-context capabilities.
RANK_REASON Release of an open-source LLM with detailed training information and benchmarks, but with restrictive licensing for commercial use. [lever_c_demoted from research: ic=1 ai=1.0]
- AMD
- HELMET
- HumanEval+
- Instella-MoE-16B-A3B
- Instinct MI300X
- Instinct MI325X
- MIT
- Moonlight-16B-A3B
- OLMo-3-7B
- OLMoE-1B-7B
- Qwen3.5-4B-Base
- ResearchRAIL
- RULER
- SmolLM3-3B-Base
- WinoGrande
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →