独立基准测试揭示了两款大型语言模型的性能指标。DBRX Instruct 在 GPQA 上获得 33.1% 的分数,在 MMLU-Pro 上获得 39.7%,在 Humanity's Last Exam 上获得 6.6%,在 LiveCodeBench 上获得 9.3%。Mistral Medium 3 表现出更高的性能,在 GPQA 上获得 57.8% 的分数,在 MMLU-Pro 上获得 76%,在 Humanity's Last Exam 上获得 4.3%,同时在 Long Context Reasoning 上显示 28% 的性能和 49.2 tokens/sec 的速度。 AI
影响 提供了 DBRX Instruct 和 Mistral Medium 3 在多个关键基准测试上的比较性能数据。
排序理由 该集群报告了两款 LLM 的独立基准测试结果,属于研究范畴。
在 Mastodon — fosstodon.org 阅读 →
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Humanity's Last Exam
- Mistral Medium 3
- MMLU-Pro
- DBRX Instruct
- LiveCodeBench
- Long Context Reasoning
AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →