PulseAugur
实时 08:36:36
English(EN) 📊 Mi:dm K 2.5 Pro — the actual numbers GPQA: 70.1% MMLU-Pro: 80.9% Humanity's Last Exam: 7.7% Long Context Reasoning: 9% Measured independently, not self-report

开源大模型在多项指标上表现强劲 · 追踪 4 个来源

根据独立测量结果,几款开源人工智能模型在各种基准测试中表现强劲。Mi:dm K 2.5 Pro 在 GPQA 上达到 70.1%,在 MMLU-Pro 上达到 80.9%;MiMo-V2-Flash 在 GPQA 上达到 83.5%,在 Humanity's Last Exam 上达到 20%;Qwen3.5 122B A10B 在 GPQA 上达到 85.7%,在 Humanity's Last Exam 上达到 23.4%;Apriel-v1.6-15B-Thinker 在 GPQA 上得分 73.3%,在 MMLU-Pro 上得分 79%。这些结果突显了开源大模型在不同评估指标上的能力正快速进步。 AI

影响 展示了开源大模型在各种基准测试中的能力取得重大进展,可能加速其采用。

排序理由 该集群报告了多款开源大模型的基准测试结果,数据来源于独立测量。

在 Mastodon — sigmoid.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 7 个来源。 我们如何撰写摘要 →

开源大模型在多项指标上表现强劲 · 追踪 4 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群报告了多款开源大模型的基准测试结果,数据来源于独立测量。
Source corroboration
7 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
45 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+3 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [7]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Mi:dm K 2.5 Pro — 实际分数 GPQA: 70.1% MMLU-Pro: 80.9% 人类终极考试: 7.7% 长上下文推理: 9% 独立测量,非自我报告

    📊 Mi:dm K 2.5 Pro — the actual numbers GPQA: 70.1% MMLU-Pro: 80.9% Humanity's Last Exam: 7.7% Long Context Reasoning: 9% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 K-EXAONE (推理) — 实际数字 GPQA: 78.3% MMLU-Pro: 83.8% 人类终极考试: 13.1% 长上下文推理: 55.7% 独立测量,非 se

    📊 K-EXAONE (Reasoning) — the actual numbers GPQA: 78.3% MMLU-Pro: 83.8% Humanity's Last Exam: 13.1% Long Context Reasoning: 55.7% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 VL 8B (推理) — 实际分数 GPQA: 57.9% MMLU-Pro: 74.9% 人类终极考试: 3.3% 长上下文推理: 31% ⚡ 128.9 tokens/秒 💰 16.1 inte

    📊 Qwen3 VL 8B (Reasoning) — the actual numbers GPQA: 57.9% MMLU-Pro: 74.9% Humanity's Last Exam: 3.3% Long Context Reasoning: 31% ⚡ 128.9 tokens/sec 💰 16.1 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LL…

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Llama 3.3 Nemotron Super 49B v1 (推理) — 实际数字 GPQA: 64.3% MMLU-Pro: 78.5% Humanity's Last Exam: 6.5% Long Context Reasoning: 17% Measured i

    📊 Llama 3.3 Nemotron Super 49B v1 (Reasoning) — the actual numbers GPQA: 64.3% MMLU-Pro: 78.5% Humanity's Last Exam: 6.5% Long Context Reasoning: 17% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  5. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 MiMo-V2-Flash (2026年2月) — 实际数字 GPQA: 83.5% 人类最后考试: 20% 长上下文推理: 64.3% SciCode: 38.3% 💰 221.3 智能点

    📊 MiMo-V2-Flash (Feb 2026) — the actual numbers GPQA: 83.5% Humanity's Last Exam: 20% Long Context Reasoning: 64.3% SciCode: 38.3% 💰 221.3 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # …

  6. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 Qwen3.5 122B A10B (推理) — 实际分数 GPQA: 85.7% 人类最后考试: 23.4% 长上下文推理: 66.7% SciCode: 42% ⚡ 147.1 tokens/秒 💰 29.

    📊 Qwen3.5 122B A10B (Reasoning) — the actual numbers GPQA: 85.7% Humanity's Last Exam: 23.4% Long Context Reasoning: 66.7% SciCode: 42% ⚡ 147.1 tokens/sec 💰 29.4 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. htm…

  7. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 Apriel-v1.6-15B-Thinker — 实际数字 GPQA: 73.3% MMLU-Pro: 79% Humanity's Last Exam: 9.8% Long Context Reasoning: 50.3% 独立测量,未进行

    📊 Apriel-v1.6-15B-Thinker — the actual numbers GPQA: 73.3% MMLU-Pro: 79% Humanity's Last Exam: 9.8% Long Context Reasoning: 50.3% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI