A South Korean startup, VIDRAFT, has developed a language model named Darwin-398B-JGOS that achieved the top rank among Korean models on the GPQA Diamond benchmark. This model, reportedly trained on approximately 24 GPUs, secured the third position globally on the benchmark, surpassing several prominent models including DeepSeek V4-Pro and NVIDIA's Nemotron 3 Ultra. The achievement is notable as Darwin-398B-JGOS was created through evolutionary merging of open models, emphasizing method over massive scale, and demonstrates strong reasoning capabilities on a benchmark designed to resist simple web lookups. AI
IMPACT Demonstrates that innovative model merging techniques can yield top-tier reasoning capabilities without massive GPU clusters.
RANK_REASON The cluster reports on a new model's performance on a specific benchmark, which falls under research.
- Darwin-398B-JGOS
- DeepSeek V4-Pro
- GPQA Diamond
- idavidrein/gpqa
- K-EXAONE-2.0-750B-A37B
- Nemotron 3 Ultra
- NVIDIA
- Qwen3.5 397B
- Solar-Open2-250B
- VIDRAFT
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →