PulseAugur
中
实时 15:56:19
English(EN) 📊 Kimi K2 0905 — the actual numbers GPQA: 76.7% MMLU-Pro: 81.9% Humanity's Last Exam: 6.4% Long Context Reasoning: 53.7% 💰 22.3 intelligence points per dollar M

LLM 基准测试显示 Kimi、Qwen3 和 Exaone 模型结果不一

独立基准测试揭示了几款大型语言模型在性能上的差异。Kimi K2 0905 在 GPQA 和 MMLU-Pro 上取得了高分,而 Qwen3 235B A22B 在这些指标上也表现良好,但在长上下文推理方面遇到困难。Exaone 4.0 在各项测试中得分较低,尤其是在推理任务方面。Qwen3 VL 4B 在其通用知识能力和处理困难推理挑战方面的表现之间存在显著差距。 AI

影响 这些基准测试突显了推理和上下文处理方面的性能差异,为未来的模型开发和选择提供了指导。

排序理由 该集群报告了多个 LLM 的基准测试结果,属于研究范畴。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →

LLM 基准测试显示 Kimi、Qwen3 和 Exaone 模型结果不一

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群报告了多个 LLM 的基准测试结果,属于研究范畴。
Source corroboration
6 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
61 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+2 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [6]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 GLM-4.6 (推理) — 实际分数 GPQA: 78% MMLU-Pro: 82.9% 人类终极考试: 14.5% 长上下文推理: 55.3% 💰 每 30.4 个智能点

    📊 GLM-4.6 (Reasoning) — the actual numbers GPQA: 78% MMLU-Pro: 82.9% Humanity's Last Exam: 14.5% Long Context Reasoning: 55.3% 💰 30.4 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Granite 4.0 1B — 实际数字 GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 4.8% Long Context Reasoning: 6% 独立测量,非自我报告

    📊 Granite 4.0 1B — the actual numbers GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 4.8% Long Context Reasoning: 6% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Kimi K2 0905 — 实际分数 GPQA: 76.7% MMLU-Pro: 81.9% 人类终极考试: 6.4% 长上下文推理: 53.7% 💰 每美元 22.3 智能点 M

    📊 Kimi K2 0905 — the actual numbers GPQA: 76.7% MMLU-Pro: 81.9% Humanity's Last Exam: 6.4% Long Context Reasoning: 53.7% 💰 22.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Exaone 4.0 1.2B(非推理)— 实际数字 GPQA: 42.4% MMLU-Pro: 50% Humanity's Last Exam: 5.7% Long Context Reasoning: 0% 独立测量,n

    📊 Exaone 4.0 1.2B (Non-reasoning) — the actual numbers GPQA: 42.4% MMLU-Pro: 50% Humanity's Last Exam: 5.7% Long Context Reasoning: 0% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  5. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Qwen3 VL 4B (Reasoning) 在 MMLU-Pro 上达到 70%,但在 Humanity's Last Exam 上仅为 4.6%——通用知识扎实与真正困难推理之间存在明显差距。https:

    Qwen3 VL 4B (Reasoning) hits 70% on MMLU-Pro but only 4.6% on Humanity's Last Exam—a clear gap between strong general knowledge and truly hard reasoning. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  6. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 235B A22B (推理) — 实际分数 GPQA: 70% MMLU-Pro: 82.8% 人类终极考试: 11% 长上下文推理: 0% 💰 每 5.1 智能点

    📊 Qwen3 235B A22B (Reasoning) — the actual numbers GPQA: 70% MMLU-Pro: 82.8% Humanity's Last Exam: 11% Long Context Reasoning: 0% 💰 5.1 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # A…