PulseAugur
EN
LIVE 12:03:03

LLM benchmarks show mixed results for Kimi, Qwen3, and Exaone models

Independent benchmarks reveal varying performance across several large language models. Kimi K2 0905 achieved strong scores on GPQA and MMLU-Pro, while Qwen3 235B A22B also performed well on these metrics but struggled with long-context reasoning. Exaone 4.0 showed lower scores across the board, particularly in reasoning tasks. Qwen3 VL 4B demonstrated a notable gap between its general knowledge capabilities and its performance on difficult reasoning challenges. AI

IMPACT These benchmarks highlight performance differences in reasoning and context handling, guiding future model development and selection.

RANK_REASON The cluster reports benchmark results for multiple LLMs, which falls under research.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 5 sources. How we write summaries →

LLM benchmarks show mixed results for Kimi, Qwen3, and Exaone models

COVERAGE [5]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Granite 4.0 1B — the actual numbers GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 4.8% Long Context Reasoning: 6% Measured independently, not self-reporte

    📊 Granite 4.0 1B — the actual numbers GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 4.8% Long Context Reasoning: 6% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Kimi K2 0905 — the actual numbers GPQA: 76.7% MMLU-Pro: 81.9% Humanity's Last Exam: 6.4% Long Context Reasoning: 53.7% 💰 22.3 intelligence points per dollar M

    📊 Kimi K2 0905 — the actual numbers GPQA: 76.7% MMLU-Pro: 81.9% Humanity's Last Exam: 6.4% Long Context Reasoning: 53.7% 💰 22.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Exaone 4.0 1.2B (Non-reasoning) — the actual numbers GPQA: 42.4% MMLU-Pro: 50% Humanity's Last Exam: 5.7% Long Context Reasoning: 0% Measured independently, n

    📊 Exaone 4.0 1.2B (Non-reasoning) — the actual numbers GPQA: 42.4% MMLU-Pro: 50% Humanity's Last Exam: 5.7% Long Context Reasoning: 0% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Qwen3 VL 4B (Reasoning) hits 70% on MMLU-Pro but only 4.6% on Humanity's Last Exam—a clear gap between strong general knowledge and truly hard reasoning. https:

    Qwen3 VL 4B (Reasoning) hits 70% on MMLU-Pro but only 4.6% on Humanity's Last Exam—a clear gap between strong general knowledge and truly hard reasoning. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  5. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 235B A22B (Reasoning) — the actual numbers GPQA: 70% MMLU-Pro: 82.8% Humanity's Last Exam: 11% Long Context Reasoning: 0% 💰 5.1 intelligence points per

    📊 Qwen3 235B A22B (Reasoning) — the actual numbers GPQA: 70% MMLU-Pro: 82.8% Humanity's Last Exam: 11% Long Context Reasoning: 0% 💰 5.1 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # A…