PulseAugur
EN
LIVE 00:03:02

DBRX Instruct and Mistral Medium 3 benchmark results revealed

Independent benchmarks reveal performance metrics for two large language models. DBRX Instruct achieved scores of 33.1% on GPQA, 39.7% on MMLU-Pro, 6.6% on Humanity's Last Exam, and 9.3% on LiveCodeBench. Mistral Medium 3 demonstrated higher performance, scoring 57.8% on GPQA, 76% on MMLU-Pro, and 4.3% on Humanity's Last Exam, while also showing 28% on Long Context Reasoning and a speed of 49.2 tokens/sec. AI

IMPACT Provides comparative performance data for DBRX Instruct and Mistral Medium 3 across several key benchmarks.

RANK_REASON The cluster reports independent benchmark results for two LLMs, which falls under research.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

DBRX Instruct and Mistral Medium 3 benchmark results revealed

COVERAGE [2]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 DBRX Instruct — the actual numbers GPQA: 33.1% MMLU-Pro: 39.7% Humanity's Last Exam: 6.6% LiveCodeBench: 9.3% Measured independently, not self-reported → http

    📊 DBRX Instruct — the actual numbers GPQA: 33.1% MMLU-Pro: 39.7% Humanity's Last Exam: 6.6% LiveCodeBench: 9.3% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Mistral Medium 3 — the actual numbers GPQA: 57.8% MMLU-Pro: 76% Humanity's Last Exam: 4.3% Long Context Reasoning: 28% ⚡ 49.2 tokens/sec 💰 15.6 intelligence p

    📊 Mistral Medium 3 — the actual numbers GPQA: 57.8% MMLU-Pro: 76% Humanity's Last Exam: 4.3% Long Context Reasoning: 28% ⚡ 49.2 tokens/sec 💰 15.6 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchm…