PulseAugur
中
实时 12:03:34
English(EN) DeepSeek R1 (Jan '25) scores 84.4% MMLU-Pro and 70.8% GPQA – but at 4.6 intelligence points per dollar, it’s the real cost story. Independently measured. https:

DeepSeek R1 模型展示了强大的基准性能,但成本效益是关键

DeepSeek 的 R1 模型于 2025 年 1 月发布,在 MMLU-Pro 和 GPQA 基准测试中取得了显著的成绩,分别达到了 84.4% 和 70.8%。然而,该模型的成本效益得到了强调,其“每美元智能点数”指标为 4.6,这表明它可能是用户的一个重要考虑因素。 AI

影响 该模型的性能和成本指标为评估模型效率的人工智能开发人员和研究人员提供了宝贵的数据。

排序理由 该条目报告了特定人工智能模型的基准分数,属于研究范畴。[lever_c_从研究降级:ic=1 ai=1.0]

在 Mastodon — mastodon.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

DeepSeek R1 模型展示了强大的基准性能,但成本效益是关键

本文如何被排名

Signal score
13 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该条目报告了特定人工智能模型的基准分数,属于研究范畴。[lever_c_从研究降级:ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    DeepSeek R1 (1月25日) MMLU-Pro得分84.4%,GPQA得分70.8%——但每美元4.6个智能点,这才是真正值得关注的成本故事。独立测量。https:

    DeepSeek R1 (Jan '25) scores 84.4% MMLU-Pro and 70.8% GPQA – but at 4.6 intelligence points per dollar, it’s the real cost story. Independently measured. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI