PulseAugur
中
实时 03:58:26
English(EN) 📊 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) — the actual numbers GPQA: 92.8% Humanity's Last Exam: 41% Long Context Reasoning: 80.3% SciCode: 51% ⚡ 81.1 toke

DeepSeek V4 Pro 和 Nova 2.0 Lite 基准测试公布 · 跟踪 2 个来源

独立基准测试揭示了两个大型语言模型 DeepSeek V4 Pro 和 Nova 2.0 Lite 的性能。DeepSeek V4 Pro 在侧重推理的任务中取得了高分,GPQA 得分为 92.8%,长上下文推理得分为 80.3%。Nova 2.0 Lite 被描述为非推理模型,在这些基准测试中的得分较低,但提供了更高的每美元智能点数指标。 AI

影响 为两个 LLM 提供比较性能数据,有助于为特定任务选择模型。

排序理由 两个 LLM 的独立基准测试结果已发布。

在 Mastodon — mastodon.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

DeepSeek V4 Pro 和 Nova 2.0 Lite 基准测试公布 · 跟踪 2 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
两个 LLM 的独立基准测试结果已发布。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
12 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 Nova 2.0 Lite (Non-reasoning) — the actual numbers GPQA: 60.3% MMLU-Pro: 74.3% Humanity's Last Exam: 2.9% Long Context Reasoning: 18.7% ⚡ 226.5 tokens/sec 💰 1

    📊 Nova 2.0 Lite (Non-reasoning) — the actual numbers GPQA: 60.3% MMLU-Pro: 74.3% Humanity's Last Exam: 2.9% Long Context Reasoning: 18.7% ⚡ 226.5 tokens/sec 💰 10.2 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM #…

  2. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 DeepSeek V4 Pro 0813 (推理,尽最大努力) — 实际数字 GPQA: 92.8% 人类最后考试: 41% 长上下文推理: 80.3% SciCode: 51% ⚡ 81.1 toke

    📊 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) — the actual numbers GPQA: 92.8% Humanity's Last Exam: 41% Long Context Reasoning: 80.3% SciCode: 51% ⚡ 81.1 tokens/sec 💰 18.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.ht…