The Darwin AI model family achieves a 90.9% score on the GPQA Diamond benchmark by using evolutionary merging of existing open-weight models, rather than traditional pretraining. This approach, which combines models like Gemma-4 and Qwen-3.5, is significantly more compute-efficient for smaller teams. While Darwin's performance places it highly on this specific benchmark, its capabilities are limited by its parent models, and its widespread adoption is largely driven by community-created copies on platforms like Hugging Face. AI
IMPACT Demonstrates a cost-effective alternative to pretraining for achieving high benchmark scores, potentially influencing future open-source model development.
RANK_REASON The item details a novel method for creating LLMs (evolutionary merging) and reports benchmark performance, fitting the research category. [lever_c_demoted from research: ic=1 ai=1.0]
- ansulev/Darwin-9B-NEG
- Apache Software License 2.0
- DeepSeek V4-Pro
- FINAL-Bench/Darwin-9B-NEG
- Gemma-4
- GLM-5.2
- GPQA Diamond
- Hugging Face
- K. E. Bartowski
- Kimi k3
- mradermacher
- Qwen-3.5
- Qwen3.5 397B
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →