StepFun has released Step-3.7-Flash, a 198 billion parameter Mixture of Experts (MoE) model that performs inference using only 11 billion active parameters per token. This architecture allows for high throughput, reaching up to 400 tokens per second, and significantly reduces computational costs. The model includes a vision encoder for image understanding and has demonstrated competitive performance against leading models like GPT-5.5 and Claude Opus-4.6 on various benchmarks, particularly in coding and visual question answering tasks. AI
IMPACT This model's cost-effective inference could accelerate the adoption of large, capable models in production environments.
RANK_REASON New model release from a frontier lab (StepFun) with detailed technical specifications and benchmark comparisons. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
- Claude Opus-4.6
- Claude Opus 4.7
- DGX Spark
- Gemini 3 Flash
- GPT-5.5
- Kimi K2.6
- llama.cpp
- Step-3.5-Flash
- Step-3.7-Flash
- StepFun
- X
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →