PulseAugur
EN
LIVE 06:49:47

LLM performance benchmarks released for Llama, GLM, and Mistral models · 4 sources tracked

Independent benchmarks reveal performance metrics for several large language models, including Llama 3.2 Instruct 90B, GLM-4.7-Flash, Mistral Large 2, and Llama 3.1 Instruct 8B. The data highlights scores across various evaluations such as GPQA, MMLU-Pro, Humanity's Last Exam, and Long Context Reasoning, with some models also reporting intelligence points per dollar. AI

IMPACT Provides comparative performance data for key LLMs, aiding developers and researchers in model selection.

RANK_REASON Independent benchmark results for multiple LLMs are presented.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 6 sources. How we write summaries →

LLM performance benchmarks released for Llama, GLM, and Mistral models · 4 sources tracked

COVERAGE [6]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Llama 3.2 Instruct 90B (Vision) — the actual numbers GPQA: 43.2% MMLU-Pro: 67.1% Humanity's Last Exam: 4.5% LiveCodeBench: 21.4% Measured independently, not s

    📊 Llama 3.2 Instruct 90B (Vision) — the actual numbers GPQA: 43.2% MMLU-Pro: 67.1% Humanity's Last Exam: 4.5% LiveCodeBench: 21.4% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 VL 4B (Reasoning) — the actual numbers GPQA: 49.4% MMLU-Pro: 70% Humanity's Last Exam: 4.6% Long Context Reasoning: 23% Measured independently, not self

    📊 Qwen3 VL 4B (Reasoning) — the actual numbers GPQA: 49.4% MMLU-Pro: 70% Humanity's Last Exam: 4.6% Long Context Reasoning: 23% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 GLM-4.7-Flash (Reasoning) — the actual numbers GPQA: 58.1% Humanity's Last Exam: 7.6% Long Context Reasoning: 40.7% SciCode: 33.7% 💰 152.3 intelligence points

    📊 GLM-4.7-Flash (Reasoning) — the actual numbers GPQA: 58.1% Humanity's Last Exam: 7.6% Long Context Reasoning: 40.7% SciCode: 33.7% 💰 152.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSourc…

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Mistral Large 2 (Nov '24) — the actual numbers GPQA: 48.6% MMLU-Pro: 69.7% Humanity's Last Exam: 3.3% Long Context Reasoning: 5.3% Measured independently, not

    📊 Mistral Large 2 (Nov '24) — the actual numbers GPQA: 48.6% MMLU-Pro: 69.7% Humanity's Last Exam: 3.3% Long Context Reasoning: 5.3% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  5. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Llama 3.1 Instruct 8B — the actual numbers GPQA: 25.9% MMLU-Pro: 47.6% Humanity's Last Exam: 5.3% Long Context Reasoning: 18% 💰 264.3 intelligence points per

    📊 Llama 3.1 Instruct 8B — the actual numbers GPQA: 25.9% MMLU-Pro: 47.6% Humanity's Last Exam: 5.3% Long Context Reasoning: 18% 💰 264.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # A…

  6. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 Qwen3 4B 2507 (Reasoning) scores 66.7% GPQA and 74.3% MMLU-Pro, but only 6.2% on Humanity's Last Exam — independent measurements reveal the gaps. Full data he

    📊 Qwen3 4B 2507 (Reasoning) scores 66.7% GPQA and 74.3% MMLU-Pro, but only 6.2% on Humanity's Last Exam — independent measurements reveal the gaps. Full data here. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI