PulseAugur
中
实时 07:39:03
English(EN) 📊 Llama 3.1 Instruct 8B — the actual numbers GPQA: 25.9% MMLU-Pro: 47.6% Humanity's Last Exam: 5.3% Long Context Reasoning: 18% 💰 264.3 intelligence points per

Llama、GLM 和 Mistral 模型发布 LLM 性能基准 · 跟踪 4 个来源

独立基准测试揭示了包括 Llama 3.2 Instruct 90B、GLM-4.7-Flash、Mistral Large 2 和 Llama 3.1 Instruct 8B 在内的多个大型语言模型的性能指标。数据突出了 GPQA、MMLU-Pro、Humanity's Last Exam 和 Long Context Reasoning 等各种评估的分数,一些模型还报告了每美元的智能点数。 AI

影响 提供关键 LLM 的比较性能数据,帮助开发人员和研究人员选择模型。

排序理由 展示了多个 LLM 的独立基准测试结果。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 7 个来源。 我们如何撰写摘要 →

Llama、GLM 和 Mistral 模型发布 LLM 性能基准 · 跟踪 4 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
展示了多个 LLM 的独立基准测试结果。
Source corroboration
7 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
53 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+3 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [7]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Llama 3.2 Instruct 90B (Vision) — 实际数字 GPQA: 43.2% MMLU-Pro: 67.1% Humanity's Last Exam: 4.5% LiveCodeBench: 21.4% 独立测量,非官方

    📊 Llama 3.2 Instruct 90B (Vision) — the actual numbers GPQA: 43.2% MMLU-Pro: 67.1% Humanity's Last Exam: 4.5% LiveCodeBench: 21.4% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Qwen3 4B 2507 (Reasoning) 在 GPQA 上达到 66.7%,在 MMLU-Pro 上达到 74.3% — 对于一个 4B 模型来说表现不错,但其在 Humanity's Last Exam 上仅 6.2% 的得分说明了推理基准的局限性

    Qwen3 4B 2507 (Reasoning) hits 66.7% on GPQA and 74.3% on MMLU-Pro — decent for a 4B model, but its 6.2% on Humanity's Last Exam shows why reasoning benchmarks still separate the small from the large. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 VL 4B (推理) — 实际分数 GPQA: 49.4% MMLU-Pro: 70% Humanity's Last Exam: 4.6% Long Context Reasoning: 23% 独立测量,非自我报告

    📊 Qwen3 VL 4B (Reasoning) — the actual numbers GPQA: 49.4% MMLU-Pro: 70% Humanity's Last Exam: 4.6% Long Context Reasoning: 23% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 GLM-4.7-Flash (推理) — 实际分数 GPQA: 58.1% 人类最后考试: 7.6% 长上下文推理: 40.7% SciCode: 33.7% 💰 152.3 智能点

    📊 GLM-4.7-Flash (Reasoning) — the actual numbers GPQA: 58.1% Humanity's Last Exam: 7.6% Long Context Reasoning: 40.7% SciCode: 33.7% 💰 152.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSourc…

  5. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Mistral Large 2 (Nov '24) — 实际数字 GPQA: 48.6% MMLU-Pro: 69.7% Humanity's Last Exam: 3.3% Long Context Reasoning: 5.3% 独立测量,非

    📊 Mistral Large 2 (Nov '24) — the actual numbers GPQA: 48.6% MMLU-Pro: 69.7% Humanity's Last Exam: 3.3% Long Context Reasoning: 5.3% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  6. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Llama 3.1 Instruct 8B — 实际数据 GPQA: 25.9% MMLU-Pro: 47.6% Humanity's Last Exam: 5.3% Long Context Reasoning: 18% 💰 每 264.3 智能点

    📊 Llama 3.1 Instruct 8B — the actual numbers GPQA: 25.9% MMLU-Pro: 47.6% Humanity's Last Exam: 5.3% Long Context Reasoning: 18% 💰 264.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # A…

  7. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 Qwen3 4B 2507(推理)在 GPQA 上得分 66.7%,在 MMLU-Pro 上得分 74.3%,但在“人类最后考试”上仅得 6.2%——独立测量揭示了差距。完整数据他

    📊 Qwen3 4B 2507 (Reasoning) scores 66.7% GPQA and 74.3% MMLU-Pro, but only 6.2% on Humanity's Last Exam — independent measurements reveal the gaps. Full data here. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI