PulseAugur
EN
LIVE 19:31:18

DeepSeek V4-Flash benchmark scores questioned due to unreleased testing harness

A recent analysis of DeepSeek's V4-Flash model reveals a significant discrepancy between its claimed performance on the Terminal-Bench 2.1 benchmark and independently verified results. DeepSeek's own published chart shows a score of 82.7, but an independent measurement by Artificial Analysis recorded 79. This 3.7-point difference is substantial, especially considering DeepSeek's claimed lead over GLM-5.2 was only 1.7 points. The discrepancy appears to stem from DeepSeek's use of an unreleased internal framework, the DeepSeek Harness, for its benchmark testing, which may inflate scores. AI

IMPACT Raises questions about the reliability of AI model benchmarks and the transparency of testing methodologies.

RANK_REASON Article analyzes and questions benchmark results published by a model developer, rather than reporting on a new release or research finding.

Read on Towards AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

DeepSeek V4-Flash benchmark scores questioned due to unreleased testing harness

COVERAGE [1]

  1. Towards AI TIER_1 English(EN) · Chew Loong Nian - AI ENGINEER ·

    DeepSeek V4-Flash vs GLM-5.2: The 1.7-Point Win Collapses When You Swap the Harness

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/deepseek-v4-flash-vs-glm-5-2-the-1-7-point-win-collapses-when-you-swap-the-harness-aa3327c87e26?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1400/1*ibqv1…