Researchers have introduced Mach-Mind-4-Flash, a 35 billion parameter Mixture-of-Experts (MoE) model that activates only 3 billion parameters. Through post-training optimization, this model achieves performance comparable to or exceeding 100 billion parameter models. The system utilizes a three-stage pipeline involving a unified RL/OPD training infrastructure, domain-specific RL experts fused via Multi-Teacher On-Policy Distillation, and Hybrid Median-length Policy Optimization for compressing reasoning chains. AI
IMPACT Demonstrates significant efficiency gains in large language models through optimized parameter activation and training techniques.
RANK_REASON Technical report detailing a new model architecture and training methodology.
- AIME'26
- arXiv
- Behavioral-SafetyBench
- BFCL-v4
- BrowseComp-zh
- ClawBench
- Hugging Face
- IFBench
- Mach-Mind-4-Flash
- mixture of experts
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →