PulseAugur
实时 23:03:11
English(EN) 📊 Qwen3 235B A22B (Reasoning) — the actual numbers GPQA: 70% MMLU-Pro: 82.8% Humanity's Last Exam: 11% Long Context Reasoning: 0% 💰 5.1 intelligence points per

Qwen3 235B A22B 模型展示基准性能,包括 GPQA 得分 70%

Qwen3 235B A22B 模型在多个基准测试中展示了性能指标,包括 GPQA 得分 70% 和 MMLU-Pro 得分 82.8%。它在人类最后考试中获得 11% 的分数,在长上下文推理任务中获得 0% 的分数。这些结果是独立测量的,表明每美元的成本效益为 5.1 个智能点。 AI

影响 提供了 Qwen3 235B A22B 模型的具体基准分数,有助于将其能力与其他大型语言模型进行比较。

排序理由 该项目报告了 AI 模型的基准测试结果,属于研究范畴。[lever_c 从研究降级:ic=1 ai=1.0]

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Qwen3 235B A22B 模型展示基准性能,包括 GPQA 得分 70%

报道来源 [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 235B A22B (Reasoning) — the actual numbers GPQA: 70% MMLU-Pro: 82.8% Humanity's Last Exam: 11% Long Context Reasoning: 0% 💰 5.1 intelligence points per

    📊 Qwen3 235B A22B (Reasoning) — the actual numbers GPQA: 70% MMLU-Pro: 82.8% Humanity's Last Exam: 11% Long Context Reasoning: 0% 💰 5.1 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # A…