Independent benchmarks for Llama 3.2 Instruct 90B (Vision) have been released, showing performance metrics across several key evaluations. The model achieved 43.2% on GPQA, 67.1% on MMLU-Pro, 4.5% on Humanity's Last Exam, and 21.4% on LiveCodeBench. These results were measured independently rather than being self-reported. AI
IMPACT Provides independent performance data for Llama 3.2 Instruct 90B, aiding in model selection and comparison.
RANK_REASON Independent benchmark results for an open-source model. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — sigmoid.social →
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Humanity's Last Exam
- LiveCodeBench
- Llama~3.2
- MMLU-Pro
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →