PulseAugur
EN
LIVE 04:16:39

Mistral Medium 3 benchmarks show 76% MMLU-Pro, 4.3% on Humanity's Last Exam

Mistral Medium 3 has been independently benchmarked, revealing its performance across several key metrics. The model achieved 57.8% on GPQA, 76% on MMLU-Pro, and 4.3% on Humanity's Last Exam. Its long context reasoning capability was measured at 28%, and it processed at a speed of 49.2 tokens per second. AI

IMPACT Provides key performance data for Mistral Medium 3, useful for model selection and comparison.

RANK_REASON Independent benchmark results for an LLM. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Mistral Medium 3 benchmarks show 76% MMLU-Pro, 4.3% on Humanity's Last Exam

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Mistral Medium 3 — the actual numbers GPQA: 57.8% MMLU-Pro: 76% Humanity's Last Exam: 4.3% Long Context Reasoning: 28% ⚡ 49.2 tokens/sec 💰 15.6 intelligence p

    📊 Mistral Medium 3 — the actual numbers GPQA: 57.8% MMLU-Pro: 76% Humanity's Last Exam: 4.3% Long Context Reasoning: 28% ⚡ 49.2 tokens/sec 💰 15.6 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchm…