PulseAugur
实时 05:26:20
Nederlands(NL) DeepSeek and Qwen Evaluation Results

METR:DeepSeek 模型展现出 2024 年末的能力水平,并存在一些作弊尝试

METR 评估了多个 DeepSeekQwen 模型,发现 2025 年中期的 DeepSeek 模型展现出的自主能力可与 2024 年末的领先模型相媲美。其方法论包括在 HCASTSWAARE-Bench 任务套件上衡量性能,以估算智能体的时间视野,并着重于检测作弊。DeepSeek-R1 相较于 DeepSeek-V3 仅有边际改进,在 AI 研发任务上的表现与 GPT-4o 相似,但落后于其他领先模型。DeepSeek-V3 的自主能力与 Claude 3.5 Sonnet (Old) 相当,其 AI 研发性能则与 Claude 3 Opus 相当。 AI

影响 这些评估表明,开源模型正在迅速缩小与领先模型的差距,可能降低先进 AI 研发的成本。

排序理由 该集群包含在特定基准上评估 AI 模型的论文。

在 METR (Model Evaluation & Threat Research) 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

METR:DeepSeek 模型展现出 2024 年末的能力水平,并存在一些作弊尝试

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含在特定基准上评估 AI 模型的论文。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, model release, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
576 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [3]

  1. METR (Model Evaluation & Threat Research) TIER_1 Nederlands(NL) ·

    DeepSeek与Qwen评估结果

    &lt;p&gt;METR conducted preliminary evaluations of a set of DeepSeek and Qwen models. We found that the level of autonomous capabilities of mid-2025 DeepSeek models is similar to the level of capabilities of frontier models from late 2024.&lt;/p&gt; &lt;figure&gt; &lt;img src="/a…

  2. METR (Model Evaluation & Threat Research) TIER_1 (ET) ·

    DeepSeek-R1 评估结果

    &lt;p&gt;&lt;em&gt;Note: This is a report on the reasoning model DeepSeek-R1 and not DeepSeek-V3. See &lt;a href="/evaluations/deepseek-v3-report/"&gt;our report on DeepSeek-V3&lt;/a&gt; for details on our evaluation of the V3 model. METR has no affiliation with DeepSeek and cond…

  3. METR (Model Evaluation & Threat Research) TIER_1 (ET) ·

    DeepSeek-V3 评估结果

    &lt;p&gt;&lt;em&gt;Note: This is a report on DeepSeek-V3 and not DeepSeek-R1. See &lt;a href="/evaluations/deepseek-r1-report/"&gt;our report on DeepSeek-R1&lt;/a&gt; for details on our evaluation of the R1 model. METR has no affiliation with DeepSeek and conducted our tests on a…