PulseAugur
EN
LIVE 13:07:51

DeepSeek V4 Pro and Nova 2.0 Lite benchmarks revealed · 2 sources tracked

Independent benchmarks reveal the performance of two large language models, DeepSeek V4 Pro and Nova 2.0 Lite. DeepSeek V4 Pro achieved high scores in reasoning-focused tasks, with 92.8% on GPQA and 80.3% on Long Context Reasoning. Nova 2.0 Lite, described as non-reasoning, showed lower scores on these benchmarks but offered a higher intelligence points per dollar metric. AI

IMPACT Provides comparative performance data for two LLMs, aiding in model selection for specific tasks.

RANK_REASON Independent benchmark results for two LLMs are published.

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

DeepSeek V4 Pro and Nova 2.0 Lite benchmarks revealed · 2 sources tracked

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Independent benchmark results for two LLMs are published.
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
9 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [2]

  1. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 Nova 2.0 Lite (Non-reasoning) — the actual numbers GPQA: 60.3% MMLU-Pro: 74.3% Humanity's Last Exam: 2.9% Long Context Reasoning: 18.7% ⚡ 226.5 tokens/sec 💰 1

    📊 Nova 2.0 Lite (Non-reasoning) — the actual numbers GPQA: 60.3% MMLU-Pro: 74.3% Humanity's Last Exam: 2.9% Long Context Reasoning: 18.7% ⚡ 226.5 tokens/sec 💰 10.2 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM #…

  2. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) — the actual numbers GPQA: 92.8% Humanity's Last Exam: 41% Long Context Reasoning: 80.3% SciCode: 51% ⚡ 81.1 toke

    📊 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) — the actual numbers GPQA: 92.8% Humanity's Last Exam: 41% Long Context Reasoning: 80.3% SciCode: 51% ⚡ 81.1 tokens/sec 💰 18.3 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.ht…