PulseAugur
实时 00:45:19
English(EN) 📊 DeepSeek-V2.5 (Dec '24) scores 76.3% on MATH-500. That’s independently measured, not self-reported—so it’s a real benchmark, not a press release. See how it s

DeepSeek-V2.5 在 MATH-500 基准上达到 76.3%

DeepSeek-V2.5 定于 2024 年 12 月发布,在 MATH-500 基准上取得了 76.3% 的分数。这一性能指标经过了独立验证,区别于新闻稿中常见的自报数据。结果可在公共排行榜上与其他模型进行比较。 AI

影响 独立验证的基准结果为模型能力提供了更清晰的图景,有助于客观比较。

排序理由 即将发布的模型的独立基准结果。 [lever_c_demoted from research: ic=1 ai=1.0]

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

DeepSeek-V2.5 在 MATH-500 基准上达到 76.3%

报道来源 [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 DeepSeek-V2.5 (12月) 在 MATH-500 上得分 76.3%。这是独立测量的结果,而非自我报告——因此是真实基准测试,而非新闻稿。看看它如何

    📊 DeepSeek-V2.5 (Dec '24) scores 76.3% on MATH-500. That’s independently measured, not self-reported—so it’s a real benchmark, not a press release. See how it stacks up against the claims: https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI