Xiaomi has developed MiMo-V2-Flash, a 309-billion-parameter Mixture-of-Experts model that leads open-source options on SWE-Bench for coding tasks. This model achieves high performance with significantly less computational cost through architectural innovations like hybrid attention and a multi-token prediction module for faster decoding. Additionally, a novel post-training technique called Multi-Teacher On-Policy Distillation allows the model to learn efficiently from multiple specialized teachers, further reducing training expenses. AI
IMPACT Sets a new standard for efficient LLM development, potentially lowering the barrier for high-performance open-source models.
RANK_REASON Xiaomi's release of a new, high-parameter LLM that tops significant benchmarks. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
- Chaintin
- Hugging Face
- llama
- MiMo-7B
- MiMo-7B-RL
- MiMo-V2.5
- MiMo-V2.5-Pro
- MiMo-V2-Flash
- mistral.ai
- MonkeyCode
- Ollama
- Xiaomi
- Xiaomi MiMo
- Mixture of Experts
- Multi-Teacher On-Policy Distillation
- Multi Token Prediction
- sliding-window attention
- SWE-Bench Multilingual
- SWE-Bench Verified
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →