Independent benchmarks reveal performance metrics for several large language models, including Llama 3.2 Instruct 90B, GLM-4.7-Flash, Mistral Large 2, and Llama 3.1 Instruct 8B. The data highlights scores across various evaluations such as GPQA, MMLU-Pro, Humanity's Last Exam, and Long Context Reasoning, with some models also reporting intelligence points per dollar. AI
IMPACT Provides comparative performance data for key LLMs, aiding developers and researchers in model selection.
RANK_REASON Independent benchmark results for multiple LLMs are presented.
Read on Mastodon — fosstodon.org →
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Humanity's Last Exam
- LiveCodeBench
- Llama~3.2
- MMLU-Pro
- Llama 3.1 Instruct 8B
- Llama 3.2 Instruct 90B (Vision)
- long-context reasoning
- GLM-4.7-Flash
- Llama 3.2 Instruct 90B
- Mistral Large 2
- SciCode
AI-generated summary · Google Gemini · from 6 sources. How we write summaries →