VIDRAFT's Darwin-180B-RSI, a 180-billion-parameter large language model, has achieved top performance on Swiss legal reasoning benchmarks, including LEXam and LEXam-hard. Notably, the model accomplished this without any domain-specific legal training, utilizing a technique called model-level recursive self-improvement (RSI). This method allows the model to identify and correct its own reasoning errors iteratively, demonstrating significant domain-agnostic reasoning capabilities. Darwin-180B-RSI reportedly outperforms models like GPT-5, Claude 4.5 Sonnet, and Gemini 2.5 Pro on these legal evaluations. AI
IMPACT Demonstrates potential for LLMs to generalize to new domains without explicit fine-tuning, impacting future model development and evaluation.
RANK_REASON Model achieves SOTA on a specialized benchmark without domain-specific training, demonstrating novel generalization capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- Claude 4.5 Sonnet
- Darwin-180B-RSI
- ETH Zurich
- Gemini 2.5 Pro
- GPT-5
- Hugging Face
- LEXam-hard
- Max Planck Institute
- University of Zurich
- VIDRAFT
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →