PulseAugur
实时 04:56:47
English(EN) 📊 GLM-5.2 (max) — the actual numbers GPQA: 89.5% Humanity's Last Exam: 40.1% Long Context Reasoning: 71.3% SciCode: 50.5% ⚡ 156.7 tokens/sec 💰 23.8 intelligence

LLM基准测试结果揭示多个模型的性能 · 追踪9个来源

一项最近的独立基准评估揭示了包括Kimi K2、Sarvam MayaNVIDIA Nemotron 3 Super 120BDeepSeek V3.2Falcon H1R-7BGLM-5.2GLM-5.1、Solar Open 100B和GLM-4.7-Flash在内的多个大型语言模型的性能指标。基准测试涵盖了推理、MMLU-Pro、人类最后考试和长上下文推理等领域,不同模型的结果差异显著。值得注意的是,GLM-5.2和DeepSeek V3.2在推理和MMLU-Pro方面表现强劲,而Kimi K2和Nemotron 3 Super 120B也展示了具有竞争力的分数。 AI

影响 为关键基准测试中的各种LLM提供了比较性能数据,帮助开发人员选择模型。

排序理由 该集群报告了多个LLM的基准测试结果,属于研究范畴。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 12 个来源。 我们如何撰写摘要 →

LLM基准测试结果揭示多个模型的性能 · 追踪9个来源

报道来源 [12]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 Max — 实际分数 GPQA: 76.4% MMLU-Pro: 84.1% 人类终极考试: 11.1% 长上下文推理: 46.7% 💰 每美元10智能点 数

    📊 Qwen3 Max — the actual numbers GPQA: 76.4% MMLU-Pro: 84.1% Humanity's Last Exam: 11.1% Long Context Reasoning: 46.7% 💰 10 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Kimi K2 — 实际数字 GPQA: 76.6% MMLU-Pro: 82.4% 人类终极考试: 7% 长上下文推理: 51% 💰 每美元 19.4 智能点 测量 i

    📊 Kimi K2 — the actual numbers GPQA: 76.6% MMLU-Pro: 82.4% Humanity's Last Exam: 7% Long Context Reasoning: 51% 💰 19.4 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Nova 2.0 Lite (medium) — 实际数据 GPQA: 76.8% MMLU-Pro: 81.3% Humanity's Last Exam: 8.6% Long Context Reasoning: 58.3% ⚡ 217.9 tokens/sec 💰 22.4 int

    📊 Nova 2.0 Lite (medium) — the actual numbers GPQA: 76.8% MMLU-Pro: 81.3% Humanity's Last Exam: 8.6% Long Context Reasoning: 58.3% ⚡ 217.9 tokens/sec 💰 22.4 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchm…

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Granite 4.0 1B — 实际数字 GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 5.1% Long Context Reasoning: 4% 独立测量,非自我报告

    📊 Granite 4.0 1B — the actual numbers GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 5.1% Long Context Reasoning: 4% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  5. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Sarvam M (Reasoning) — 实际分数 GPQA: 41.6% MMLU-Pro: 69.6% Humanity's Last Exam: 3.3% Long Context Reasoning: 0% 独立测量,非自我报告

    📊 Sarvam M (Reasoning) — the actual numbers GPQA: 41.6% MMLU-Pro: 69.6% Humanity's Last Exam: 3.3% Long Context Reasoning: 0% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  6. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 NVIDIA Nemotron 3 Super 120B A12B (推理) — 实际数字 GPQA: 80% 人类最后考试: 19.2% 长上下文推理: 60% SciCode: 36% ⚡ 157.5 toke

    📊 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) — the actual numbers GPQA: 80% Humanity's Last Exam: 19.2% Long Context Reasoning: 60% SciCode: 36% ⚡ 157.5 tokens/sec 💰 66.7 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/lead…

  7. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 DeepSeek V3.2 (Reasoning) — 实际分数 GPQA: 84% MMLU-Pro: 86.2% Humanity's Last Exam: 22.2% Long Context Reasoning: 65% 💰 101.6 intelligence points

    📊 DeepSeek V3.2 (Reasoning) — the actual numbers GPQA: 84% MMLU-Pro: 86.2% Humanity's Last Exam: 22.2% Long Context Reasoning: 65% 💰 101.6 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # …

  8. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Falcon-H1R-7B — 实际分数 GPQA: 66.1% MMLU-Pro: 72.5% Humanity's Last Exam: 10.8% Long Context Reasoning: 8.7% 独立测量,非自我报告

    📊 Falcon-H1R-7B — the actual numbers GPQA: 66.1% MMLU-Pro: 72.5% Humanity's Last Exam: 10.8% Long Context Reasoning: 8.7% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  9. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 GLM-5.2 (max) — 实际分数 GPQA: 89.5% Humanity's Last Exam: 40.1% Long Context Reasoning: 71.3% SciCode: 50.5% ⚡ 156.7 tokens/秒 💰 23.8 intelligence

    📊 GLM-5.2 (max) — the actual numbers GPQA: 89.5% Humanity's Last Exam: 40.1% Long Context Reasoning: 71.3% SciCode: 50.5% ⚡ 156.7 tokens/sec 💰 23.8 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benc…

  10. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 GLM-5.1 (推理) — 实际分数 GPQA: 86.8% 人类最后考试: 28% 长上下文推理: 62.3% SciCode: 43.8% 💰 每美元 18.8 智能点

    📊 GLM-5.1 (Reasoning) — the actual numbers GPQA: 86.8% Humanity's Last Exam: 28% Long Context Reasoning: 62.3% SciCode: 43.8% 💰 18.8 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  11. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 Solar Open 100B (推理) — 实际分数 GPQA: 65.7% 人类最后考试: 9.2% 长上下文推理: 36% SciCode: 26.9% 独立测量,非

    📊 Solar Open 100B (Reasoning) — the actual numbers GPQA: 65.7% Humanity's Last Exam: 9.2% Long Context Reasoning: 36% SciCode: 26.9% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  12. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 GLM-4.7-Flash (非推理能力) — 实际分数 GPQA: 45.2% 人类最后考试: 4.9% 长上下文推理: 14.7% SciCode: 25.5% ⚡ 179.6 tokens/秒 💰 10

    📊 GLM-4.7-Flash (Non-reasoning) — the actual numbers GPQA: 45.2% Humanity's Last Exam: 4.9% Long Context Reasoning: 14.7% SciCode: 25.5% ⚡ 179.6 tokens/sec 💰 101.3 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. h…