Independent benchmarks reveal performance metrics for two large language models. DBRX Instruct achieved scores of 33.1% on GPQA, 39.7% on MMLU-Pro, 6.6% on Humanity's Last Exam, and 9.3% on LiveCodeBench. Mistral Medium 3 demonstrated higher performance, scoring 57.8% on GPQA, 76% on MMLU-Pro, and 4.3% on Humanity's Last Exam, while also showing 28% on Long Context Reasoning and a speed of 49.2 tokens/sec. AI
IMPACT Provides comparative performance data for DBRX Instruct and Mistral Medium 3 across several key benchmarks.
RANK_REASON The cluster reports independent benchmark results for two LLMs, which falls under research.
Read on Mastodon — fosstodon.org →
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Humanity's Last Exam
- Mistral Medium 3
- MMLU-Pro
- DBRX Instruct
- LiveCodeBench
- Long Context Reasoning
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →