PulseAugur
EN
LIVE 14:35:19

Llama 3.2 Instruct 90B (Vision) benchmarks released

Independent benchmarks for Llama 3.2 Instruct 90B (Vision) have been released, showing performance metrics across several key evaluations. The model achieved 43.2% on GPQA, 67.1% on MMLU-Pro, 4.5% on Humanity's Last Exam, and 21.4% on LiveCodeBench. These results were measured independently rather than being self-reported. AI

IMPACT Provides independent performance data for Llama 3.2 Instruct 90B, aiding in model selection and comparison.

RANK_REASON Independent benchmark results for an open-source model. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — sigmoid.social →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Llama 3.2 Instruct 90B (Vision) benchmarks released

COVERAGE [1]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Llama 3.2 Instruct 90B (Vision) — the actual numbers GPQA: 43.2% MMLU-Pro: 67.1% Humanity's Last Exam: 4.5% LiveCodeBench: 21.4% Measured independently, not s

    📊 Llama 3.2 Instruct 90B (Vision) — the actual numbers GPQA: 43.2% MMLU-Pro: 67.1% Humanity's Last Exam: 4.5% LiveCodeBench: 21.4% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI