PulseAugur
实时 14:15:03
English(EN) 📊 Kimi K2 0905 — the actual numbers GPQA: 76.7% MMLU-Pro: 81.9% Humanity's Last Exam: 6.4% Long Context Reasoning: 53.7% 💰 22.3 intelligence points per dollar M

LLM 基准测试显示 Kimi、Qwen3 和 Exaone 模型结果不一

独立基准测试揭示了几款大型语言模型在性能上的差异。Kimi K2 0905 在 GPQA 和 MMLU-Pro 上取得了高分,而 Qwen3 235B A22B 在这些指标上也表现良好,但在长上下文推理方面遇到困难。Exaone 4.0 在各项测试中得分较低,尤其是在推理任务方面。Qwen3 VL 4B 在其通用知识能力和处理困难推理挑战方面的表现之间存在显著差距。 AI

影响 这些基准测试突显了推理和上下文处理方面的性能差异,为未来的模型开发和选择提供了指导。

排序理由 该集群报告了多个 LLM 的基准测试结果,属于研究范畴。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

LLM 基准测试显示 Kimi、Qwen3 和 Exaone 模型结果不一

报道来源 [5]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Granite 4.0 1B — 实际数字 GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 4.8% Long Context Reasoning: 6% 独立测量,非自我报告

    📊 Granite 4.0 1B — the actual numbers GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 4.8% Long Context Reasoning: 6% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Kimi K2 0905 — 实际分数 GPQA: 76.7% MMLU-Pro: 81.9% 人类终极考试: 6.4% 长上下文推理: 53.7% 💰 每美元 22.3 智能点 M

    📊 Kimi K2 0905 — the actual numbers GPQA: 76.7% MMLU-Pro: 81.9% Humanity's Last Exam: 6.4% Long Context Reasoning: 53.7% 💰 22.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Exaone 4.0 1.2B(非推理)— 实际数字 GPQA: 42.4% MMLU-Pro: 50% Humanity's Last Exam: 5.7% Long Context Reasoning: 0% 独立测量,n

    📊 Exaone 4.0 1.2B (Non-reasoning) — the actual numbers GPQA: 42.4% MMLU-Pro: 50% Humanity's Last Exam: 5.7% Long Context Reasoning: 0% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Qwen3 VL 4B (Reasoning) 在 MMLU-Pro 上达到 70%,但在 Humanity's Last Exam 上仅为 4.6%——通用知识扎实与真正困难推理之间存在明显差距。https:

    Qwen3 VL 4B (Reasoning) hits 70% on MMLU-Pro but only 4.6% on Humanity's Last Exam—a clear gap between strong general knowledge and truly hard reasoning. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  5. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 235B A22B (推理) — 实际分数 GPQA: 70% MMLU-Pro: 82.8% 人类终极考试: 11% 长上下文推理: 0% 💰 每 5.1 智能点

    📊 Qwen3 235B A22B (Reasoning) — the actual numbers GPQA: 70% MMLU-Pro: 82.8% Humanity's Last Exam: 11% Long Context Reasoning: 0% 💰 5.1 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # A…