PulseAugur
实时 05:12:13
English(EN) 📊 Llama 3.1 Instruct 8B — the actual numbers GPQA: 25.9% MMLU-Pro: 47.6% Humanity's Last Exam: 5.3% Long Context Reasoning: 18% 💰 264.3 intelligence points per

Llama、GLM 和 Mistral 模型发布 LLM 性能基准 · 跟踪 4 个来源

独立基准测试揭示了包括 Llama 3.2 Instruct 90BGLM-4.7-FlashMistral Large 2Llama 3.1 Instruct 8B 在内的多个大型语言模型的性能指标。数据突出了 GPQA、MMLU-ProHumanity's Last Exam 和 Long Context Reasoning 等各种评估的分数,一些模型还报告了每美元的智能点数。 AI

影响 提供关键 LLM 的比较性能数据,帮助开发人员和研究人员选择模型。

排序理由 展示了多个 LLM 的独立基准测试结果。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →

Llama、GLM 和 Mistral 模型发布 LLM 性能基准 · 跟踪 4 个来源

报道来源 [6]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Llama 3.2 Instruct 90B (Vision) — 实际数字 GPQA: 43.2% MMLU-Pro: 67.1% Humanity's Last Exam: 4.5% LiveCodeBench: 21.4% 独立测量,非官方

    📊 Llama 3.2 Instruct 90B (Vision) — the actual numbers GPQA: 43.2% MMLU-Pro: 67.1% Humanity's Last Exam: 4.5% LiveCodeBench: 21.4% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 VL 4B (推理) — 实际分数 GPQA: 49.4% MMLU-Pro: 70% Humanity's Last Exam: 4.6% Long Context Reasoning: 23% 独立测量,非自我报告

    📊 Qwen3 VL 4B (Reasoning) — the actual numbers GPQA: 49.4% MMLU-Pro: 70% Humanity's Last Exam: 4.6% Long Context Reasoning: 23% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 GLM-4.7-Flash (推理) — 实际分数 GPQA: 58.1% 人类最后考试: 7.6% 长上下文推理: 40.7% SciCode: 33.7% 💰 152.3 智能点

    📊 GLM-4.7-Flash (Reasoning) — the actual numbers GPQA: 58.1% Humanity's Last Exam: 7.6% Long Context Reasoning: 40.7% SciCode: 33.7% 💰 152.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSourc…

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Mistral Large 2 (Nov '24) — 实际数字 GPQA: 48.6% MMLU-Pro: 69.7% Humanity's Last Exam: 3.3% Long Context Reasoning: 5.3% 独立测量,非

    📊 Mistral Large 2 (Nov '24) — the actual numbers GPQA: 48.6% MMLU-Pro: 69.7% Humanity's Last Exam: 3.3% Long Context Reasoning: 5.3% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  5. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Llama 3.1 Instruct 8B — 实际数据 GPQA: 25.9% MMLU-Pro: 47.6% Humanity's Last Exam: 5.3% Long Context Reasoning: 18% 💰 每 264.3 智能点

    📊 Llama 3.1 Instruct 8B — the actual numbers GPQA: 25.9% MMLU-Pro: 47.6% Humanity's Last Exam: 5.3% Long Context Reasoning: 18% 💰 264.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # A…

  6. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 Qwen3 4B 2507(推理)在 GPQA 上得分 66.7%,在 MMLU-Pro 上得分 74.3%,但在“人类最后考试”上仅得 6.2%——独立测量揭示了差距。完整数据他

    📊 Qwen3 4B 2507 (Reasoning) scores 66.7% GPQA and 74.3% MMLU-Pro, but only 6.2% on Humanity's Last Exam — independent measurements reveal the gaps. Full data here. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI