PulseAugur
中
实时 21:23:47
English(EN) 📊 DeepSeek V4 Pro 0813 (Non-reasoning) — the actual numbers Humanity's Last Exam: 10.6% Long Context Reasoning: 50.7% SciCode: 39.9% ⚡ 160.7 tokens/sec 💰 10.3 i

DeepSeek V4 Pro 0813 基准测试显示 Humanity's Last Exam 得分为 10.6%

DeepSeek V4 Pro 0813 在包括 Humanity's Last Exam、Long Context Reasoning 和 SciCode 在内的基准测试中展示了具体的性能指标。该模型在 Humanity's Last Exam 上达到 10.6%,在 Long Context Reasoning 上达到 50.7%,在 SciCode 上达到 39.9%。此外,它每秒处理 160.7 个 token,每美元达到 10.3 智能点,这些结果均经过独立测量。 AI

影响 提供了 DeepSeek V4 Pro 0813 模型在关键基准测试上的具体性能数据,有助于比较 AI 能力。

排序理由 该项目报告了 AI 模型的基准测试结果,属于研究范畴。[lever_c_research降级:ic=1 ai=1.0]

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

DeepSeek V4 Pro 0813 基准测试显示 Humanity's Last Exam 得分为 10.6%

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目报告了 AI 模型的基准测试结果,属于研究范畴。[lever_c_research降级:ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 DeepSeek V4 Pro 0813(非推理)— 实际数字 人类终极考试:10.6% 长上下文推理:50.7% SciCode:39.9% ⚡ 160.7 tokens/秒 💰 10.3 i

    📊 DeepSeek V4 Pro 0813 (Non-reasoning) — the actual numbers Humanity's Last Exam: 10.6% Long Context Reasoning: 50.7% SciCode: 39.9% ⚡ 160.7 tokens/sec 💰 10.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benc…