An open-weight model named Darwin-180B-RSI has achieved the top position on 10 official Hugging Face leaderboards, surpassing all other participating organizations. This model, built by VIDRAFT on Alibaba's Qwen3.8-Flash-Next, demonstrates strong performance not only in academic benchmarks like MMLU-Pro and GPQA but also in practical, enterprise-focused tasks such as document parsing (MDPBench), structured data output (IFStruct), and field extraction from PDFs (ExtractBench). The model's training methodology, known as recursive self-improvement (RSI), allows it to refine its capabilities by solving problems, verifying its own solutions, and retraining on confirmed correct outputs without human-labeled data. AI
IMPACT Sets a new standard for open-weight models in both academic and practical enterprise tasks, potentially accelerating adoption of advanced AI agents.
RANK_REASON Open-weight model release with multiple benchmark wins and details on training methodology. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
- Alibaba Group
- Darwin-180B-RSI
- ExtractBench
- Gate Arcade
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Hugging Face
- IFStruct
- Liquid Ai
- MDPBench
- MMLU-Pro
- Qwen3.8-Flash-Next
- VIDRAFT
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →