A viral debate on China's Zhihu platform scrutinized a wave of Korean AI models that claimed to outperform DeepSeek. The critique focused on several points: models allegedly building upon open-weight models from other labs (like SK Telecom's A.X K2 using Alibaba's Qwen2.5), narrow benchmark selection that ignored broader performance metrics, and claims timed to secure government funding. Amidst these criticisms, the startup VIDRAFT emerged as an outlier, with its Darwin-398B-JGOS model achieving a specific, verifiable ranking on the GPQA benchmark, demonstrating the value of precise and reproducible claims. AI
IMPACT Highlights the importance of transparent methodology and verifiable claims in AI model development and benchmarking.
RANK_REASON The item discusses a debate and critique of AI models rather than a direct release or product launch.
- Alibaba Group
- A.X K2
- Darwin-398B-JGOS
- DeepSeek
- DeepSeek V4-Pro
- GPQA
- K-EXAONE 2.0
- LG Group
- Qwen2.5
- SK Telecom
- Solar Open 2
- TechWalker
- Tencent News
- Upstage
- VIDRAFT
- Zhihu
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →