Several open-source AI models have demonstrated strong performance on various benchmarks, according to independent measurements. Mi:dm K 2.5 Pro achieved 70.1% on GPQA and 80.9% on MMLU-Pro, while MiMo-V2-Flash showed 83.5% on GPQA and 20% on Humanity's Last Exam. Qwen3.5 122B A10B reached 85.7% on GPQA and 23.4% on Humanity's Last Exam, and Apriel-v1.6-15B-Thinker scored 73.3% on GPQA and 79% on MMLU-Pro. These results highlight the rapid progress in open-source LLM capabilities across different evaluation metrics. AI
IMPACT Demonstrates significant advancements in open-source LLM capabilities across various benchmarks, potentially accelerating adoption.
RANK_REASON Cluster reports benchmark results for multiple open-source LLMs, sourced from independent measurements.
Read on Mastodon — sigmoid.social →
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Humanity's Last Exam
- long-context reasoning
- MiMo-V2-Flash
- SciCode
- Apriel-v1.6-15B-Thinker
- MMLU-Pro
- Mi:dm K 2.5 Pro
- Qwen3.5 122B A10B
AI-generated summary · Google Gemini · from 7 sources. How we write summaries →