Mistral Medium 3 has demonstrated performance across several benchmarks, achieving 57.8% on GPQA and 76% on MMLU-Pro. The model also scored 4.1% on Humanity's Last Exam and 31.7% in Long Context Reasoning. Notably, it offers 15.6 intelligence points per dollar, indicating a cost-effective performance. AI
IMPACT Provides performance metrics for Mistral Medium 3, useful for comparing AI model capabilities and cost-effectiveness.
RANK_REASON The item reports benchmark results for an AI model, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Humanity's Last Exam
- long-context reasoning
- Mistral Medium 3
- MMLU-Pro
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →