PulseAugur
EN
LIVE 11:21:51

StepFun's 198B MoE model offers pro-tier performance at low cost

StepFun has released Step-3.7-Flash, a 198 billion parameter Mixture of Experts (MoE) model that performs inference using only 11 billion active parameters per token. This architecture allows for high throughput, reaching up to 400 tokens per second, and significantly reduces computational costs. The model includes a vision encoder for image understanding and has demonstrated competitive performance against leading models like GPT-5.5 and Claude Opus-4.6 on various benchmarks, particularly in coding and visual question answering tasks. AI

IMPACT This model's cost-effective inference could accelerate the adoption of large, capable models in production environments.

RANK_REASON New model release from a frontier lab (StepFun) with detailed technical specifications and benchmark comparisons. [lever_c_demoted from frontier_release: ic=1 ai=1.0]

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

StepFun's 198B MoE model offers pro-tier performance at low cost

How we ranked this

Signal score
45 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
New model release from a frontier lab (StepFun) with detailed technical specifications and benchmark comparisons. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Dishant Sharma ·

    Step 3.7 Flash: The 198B MoE Model Everyone Is Actually Running

    <p>someone on X posted a photo of a DGX Spark sitting on a regular desk with a terminal window running step 3.7 flash. a 198 billion parameter vision model. on a box that fits next to a monitor. and it was not a flex post. it was a "here is the config that saved me three hours" p…