Mistral Medium 3 has been independently benchmarked, revealing its performance across several key metrics. The model achieved 57.8% on GPQA, 76% on MMLU-Pro, and 4.3% on Humanity's Last Exam. Its long context reasoning capability was measured at 28%, and it processed at a speed of 49.2 tokens per second. AI
IMPACT Provides key performance data for Mistral Medium 3, useful for model selection and comparison.
RANK_REASON Independent benchmark results for an LLM. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →