DeepSeek is retiring its flagship DeepSeek V4-Pro model, rerouting all requests to its V4.1-Flash variant. This decision follows internal and external testing that indicated V4.1-Flash outperforms V4-Pro in capability, cost, and speed, even surpassing models like Claude Opus 5 and GPT 5.6 "Sol" on benchmarks like Terminal-Bench 2.1. However, the V4.1-Flash model shows a regression in factual question answering, with SimpleQA scores dropping significantly, suggesting a trade-off between agentic performance and world knowledge. AI
IMPACT This move signals a potential shift in the industry towards prioritizing cost-efficiency and speed, even at the expense of some factual recall, influencing future model development and pricing strategies.
RANK_REASON DeepSeek is a frontier lab, and this is a significant change to its flagship model offering, effectively retiring one and promoting another.
Read on Mastodon — mastodon.social →
- Claude Opus 5
- DeepSeek
- DeepSeek V4-Pro
- GLM 5.3
- GPT 5.6 "Sol"
- Hermes Agent
- K2.7 Code
- Kimi k3
- Moonshot
- Opus 4.8
- SimpleQA
- Terminal-Bench 2.1
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →