PulseAugur
实时 14:35:32
English(EN) 📊 Llama 3.2 Instruct 90B (Vision) — the actual numbers GPQA: 43.2% MMLU-Pro: 67.1% Humanity's Last Exam: 4.5% LiveCodeBench: 21.4% Measured independently, not s

Llama 3.2 Instruct 90B (Vision) 基准测试发布

Llama 3.2 Instruct 90B (Vision) 的独立基准测试已发布,展示了其在多项关键评估中的性能指标。该模型在 GPQA 上 đạt 43.2%,在 MMLU-Pro 上 đạt 67.1%,在 Humanity's Last Exam 上 đạt 4.5%,在 LiveCodeBench 上 đạt 21.4%。这些结果是独立测量的,而非自我报告。 AI

影响 提供了 Llama 3.2 Instruct 90B 的独立性能数据,有助于模型选择和比较。

排序理由 开源模型的独立基准测试结果。 [lever_c_demoted from research: ic=1 ai=1.0]

在 Mastodon — sigmoid.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Llama 3.2 Instruct 90B (Vision) 基准测试发布

报道来源 [1]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Llama 3.2 Instruct 90B (Vision) — 实际数字 GPQA: 43.2% MMLU-Pro: 67.1% Humanity's Last Exam: 4.5% LiveCodeBench: 21.4% 独立测量,非官方

    📊 Llama 3.2 Instruct 90B (Vision) — the actual numbers GPQA: 43.2% MMLU-Pro: 67.1% Humanity's Last Exam: 4.5% LiveCodeBench: 21.4% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI