PulseAugur
EN
LIVE 09:55:52

LLM performance benchmarks released for Llama, GLM, and Mistral models · 4 sources tracked

Independent benchmarks reveal performance metrics for several large language models, including Llama 3.2 Instruct 90B, GLM-4.7-Flash, Mistral Large 2, and Llama 3.1 Instruct 8B. The data highlights scores across various evaluations such as GPQA, MMLU-Pro, Humanity's Last Exam, and Long Context Reasoning, with some models also reporting intelligence points per dollar. AI

IMPACT Provides comparative performance data for key LLMs, aiding developers and researchers in model selection.

RANK_REASON Independent benchmark results for multiple LLMs are presented.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 7 sources. How we write summaries →

LLM performance benchmarks released for Llama, GLM, and Mistral models · 4 sources tracked

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Independent benchmark results for multiple LLMs are presented.
Source corroboration
7 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
53 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+3 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [7]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Llama 3.2 Instruct 90B (Vision) — the actual numbers GPQA: 43.2% MMLU-Pro: 67.1% Humanity's Last Exam: 4.5% LiveCodeBench: 21.4% Measured independently, not s

    📊 Llama 3.2 Instruct 90B (Vision) — the actual numbers GPQA: 43.2% MMLU-Pro: 67.1% Humanity's Last Exam: 4.5% LiveCodeBench: 21.4% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Qwen3 4B 2507 (Reasoning) hits 66.7% on GPQA and 74.3% on MMLU-Pro — decent for a 4B model, but its 6.2% on Humanity's Last Exam shows why reasoning benchmarks

    Qwen3 4B 2507 (Reasoning) hits 66.7% on GPQA and 74.3% on MMLU-Pro — decent for a 4B model, but its 6.2% on Humanity's Last Exam shows why reasoning benchmarks still separate the small from the large. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 VL 4B (Reasoning) — the actual numbers GPQA: 49.4% MMLU-Pro: 70% Humanity's Last Exam: 4.6% Long Context Reasoning: 23% Measured independently, not self

    📊 Qwen3 VL 4B (Reasoning) — the actual numbers GPQA: 49.4% MMLU-Pro: 70% Humanity's Last Exam: 4.6% Long Context Reasoning: 23% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 GLM-4.7-Flash (Reasoning) — the actual numbers GPQA: 58.1% Humanity's Last Exam: 7.6% Long Context Reasoning: 40.7% SciCode: 33.7% 💰 152.3 intelligence points

    📊 GLM-4.7-Flash (Reasoning) — the actual numbers GPQA: 58.1% Humanity's Last Exam: 7.6% Long Context Reasoning: 40.7% SciCode: 33.7% 💰 152.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSourc…

  5. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Mistral Large 2 (Nov '24) — the actual numbers GPQA: 48.6% MMLU-Pro: 69.7% Humanity's Last Exam: 3.3% Long Context Reasoning: 5.3% Measured independently, not

    📊 Mistral Large 2 (Nov '24) — the actual numbers GPQA: 48.6% MMLU-Pro: 69.7% Humanity's Last Exam: 3.3% Long Context Reasoning: 5.3% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  6. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Llama 3.1 Instruct 8B — the actual numbers GPQA: 25.9% MMLU-Pro: 47.6% Humanity's Last Exam: 5.3% Long Context Reasoning: 18% 💰 264.3 intelligence points per

    📊 Llama 3.1 Instruct 8B — the actual numbers GPQA: 25.9% MMLU-Pro: 47.6% Humanity's Last Exam: 5.3% Long Context Reasoning: 18% 💰 264.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # A…

  7. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 Qwen3 4B 2507 (Reasoning) scores 66.7% GPQA and 74.3% MMLU-Pro, but only 6.2% on Humanity's Last Exam — independent measurements reveal the gaps. Full data he

    📊 Qwen3 4B 2507 (Reasoning) scores 66.7% GPQA and 74.3% MMLU-Pro, but only 6.2% on Humanity's Last Exam — independent measurements reveal the gaps. Full data here. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI