独立基准测试揭示了包括 Llama 3.2 Instruct 90B、GLM-4.7-Flash、Mistral Large 2 和 Llama 3.1 Instruct 8B 在内的多个大型语言模型的性能指标。数据突出了 GPQA、MMLU-Pro、Humanity's Last Exam 和 Long Context Reasoning 等各种评估的分数,一些模型还报告了每美元的智能点数。 AI
影响 提供关键 LLM 的比较性能数据,帮助开发人员和研究人员选择模型。
排序理由 展示了多个 LLM 的独立基准测试结果。
在 Mastodon — fosstodon.org 阅读 →
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Humanity's Last Exam
- LiveCodeBench
- Llama~3.2
- MMLU-Pro
- Llama 3.1 Instruct 8B
- Llama 3.2 Instruct 90B (Vision)
- long-context reasoning
- GLM-4.7-Flash
- Llama 3.2 Instruct 90B
- Mistral Large 2
- SciCode
AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →