PulseAugur
EN
LIVE 14:06:46

LLM benchmarks show mixed results for Kimi, Qwen3, and Exaone models

Independent benchmarks reveal varying performance across several large language models. Kimi K2 0905 achieved strong scores on GPQA and MMLU-Pro, while Qwen3 235B A22B also performed well on these metrics but struggled with long-context reasoning. Exaone 4.0 showed lower scores across the board, particularly in reasoning tasks. Qwen3 VL 4B demonstrated a notable gap between its general knowledge capabilities and its performance on difficult reasoning challenges. AI

IMPACT These benchmarks highlight performance differences in reasoning and context handling, guiding future model development and selection.

RANK_REASON The cluster reports benchmark results for multiple LLMs, which falls under research.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 6 sources. How we write summaries →

LLM benchmarks show mixed results for Kimi, Qwen3, and Exaone models

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
The cluster reports benchmark results for multiple LLMs, which falls under research.
Source corroboration
6 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
50 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+2 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [6]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 GLM-4.6 (Reasoning) — the actual numbers GPQA: 78% MMLU-Pro: 82.9% Humanity's Last Exam: 14.5% Long Context Reasoning: 55.3% 💰 30.4 intelligence points per do

    📊 GLM-4.6 (Reasoning) — the actual numbers GPQA: 78% MMLU-Pro: 82.9% Humanity's Last Exam: 14.5% Long Context Reasoning: 55.3% 💰 30.4 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Granite 4.0 1B — the actual numbers GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 4.8% Long Context Reasoning: 6% Measured independently, not self-reporte

    📊 Granite 4.0 1B — the actual numbers GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 4.8% Long Context Reasoning: 6% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Kimi K2 0905 — the actual numbers GPQA: 76.7% MMLU-Pro: 81.9% Humanity's Last Exam: 6.4% Long Context Reasoning: 53.7% 💰 22.3 intelligence points per dollar M

    📊 Kimi K2 0905 — the actual numbers GPQA: 76.7% MMLU-Pro: 81.9% Humanity's Last Exam: 6.4% Long Context Reasoning: 53.7% 💰 22.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Exaone 4.0 1.2B (Non-reasoning) — the actual numbers GPQA: 42.4% MMLU-Pro: 50% Humanity's Last Exam: 5.7% Long Context Reasoning: 0% Measured independently, n

    📊 Exaone 4.0 1.2B (Non-reasoning) — the actual numbers GPQA: 42.4% MMLU-Pro: 50% Humanity's Last Exam: 5.7% Long Context Reasoning: 0% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  5. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Qwen3 VL 4B (Reasoning) hits 70% on MMLU-Pro but only 4.6% on Humanity's Last Exam—a clear gap between strong general knowledge and truly hard reasoning. https:

    Qwen3 VL 4B (Reasoning) hits 70% on MMLU-Pro but only 4.6% on Humanity's Last Exam—a clear gap between strong general knowledge and truly hard reasoning. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  6. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 235B A22B (Reasoning) — the actual numbers GPQA: 70% MMLU-Pro: 82.8% Humanity's Last Exam: 11% Long Context Reasoning: 0% 💰 5.1 intelligence points per

    📊 Qwen3 235B A22B (Reasoning) — the actual numbers GPQA: 70% MMLU-Pro: 82.8% Humanity's Last Exam: 11% Long Context Reasoning: 0% 💰 5.1 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # A…