Open-weight Mixture of Experts (MoE) models in 2026 are diverging significantly in their total parameter count versus active parameter count. This means the total weights, which dictate memory requirements and loading times, are much larger than the active weights, which determine the computational cost per token. For instance, Kimi K3 boasts 2.8 trillion total parameters but only activates around 104 billion, while DeepSeek's V4.1-Flash has 552 billion total parameters but activates only 8 billion. This trend allows models to store vast amounts of knowledge while keeping per-token compute costs manageable, though it complicates budgeting and deployment by creating separate memory and compute budgets. AI
IMPACT This architectural shift in MoE models necessitates new strategies for budgeting and deploying AI systems, as memory and compute costs are now decoupled.
RANK_REASON Discusses trends in open-weight model architecture and performance metrics. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →