PulseAugur
实时 02:30:03
English(EN) 📊 Mistral Medium 3 — the actual numbers GPQA: 57.8% MMLU-Pro: 76% Humanity's Last Exam: 4.3% Long Context Reasoning: 28% ⚡ 49.2 tokens/sec 💰 15.6 intelligence p

DBRX Instruct 和 Mistral Medium 3 的基准测试结果公布

独立基准测试揭示了两款大型语言模型的性能指标。DBRX Instruct 在 GPQA 上获得 33.1% 的分数,在 MMLU-Pro 上获得 39.7%,在 Humanity's Last Exam 上获得 6.6%,在 LiveCodeBench 上获得 9.3%。Mistral Medium 3 表现出更高的性能,在 GPQA 上获得 57.8% 的分数,在 MMLU-Pro 上获得 76%,在 Humanity's Last Exam 上获得 4.3%,同时在 Long Context Reasoning 上显示 28% 的性能和 49.2 tokens/sec 的速度。 AI

影响 提供了 DBRX InstructMistral Medium 3 在多个关键基准测试上的比较性能数据。

排序理由 该集群报告了两款 LLM 的独立基准测试结果,属于研究范畴。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

DBRX Instruct 和 Mistral Medium 3 的基准测试结果公布

报道来源 [2]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 DBRX Instruct — 实际数字 GPQA: 33.1% MMLU-Pro: 39.7% Humanity's Last Exam: 6.6% LiveCodeBench: 9.3% 独立测量,非自我报告 → http

    📊 DBRX Instruct — the actual numbers GPQA: 33.1% MMLU-Pro: 39.7% Humanity's Last Exam: 6.6% LiveCodeBench: 9.3% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Mistral Medium 3 — 实际数字 GPQA: 57.8% MMLU-Pro: 76% Humanity's Last Exam: 4.3% Long Context Reasoning: 28% ⚡ 49.2 tokens/秒 💰 15.6 intelligence p

    📊 Mistral Medium 3 — the actual numbers GPQA: 57.8% MMLU-Pro: 76% Humanity's Last Exam: 4.3% Long Context Reasoning: 28% ⚡ 49.2 tokens/sec 💰 15.6 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchm…