PulseAugur
中
实时 03:29:27
Deutsch(DE) RT @jun_song: Ich habe DeepSeek-V4-Pro-0813 und Grok-4.6 in meiner Agent-Framework getestet. Die Leistung enttäuscht etwas im Vergleich zu den Benchmarks (getes

DeepSeek-V4-Pro-0813 和 Grok-4.6 在代理框架中的性能测试 · 跟踪 4 个来源

一位用户在代理框架中测试了 DeepSeek-V4-Pro-0813 和 Grok-4.6,发现它们的性能与基准测试相比略有令人失望。尽管如此,据报道,这两种模型仍然优于 Opus-5,而 Opus-5 目前受到计算资源的限制。另一位用户指出 DeepSeek 出色的推理能力,实现了高缓存率。 AI

影响 提供用户层面的见解,了解领先的 AI 模型在代理框架中的实际性能,突显基准测试与实际应用之间可能存在的差异。

排序理由 用户对模型性能的测试和评论,并非官方发布或基准测试。

在 Mastodon — sigmoid.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

DeepSeek-V4-Pro-0813 和 Grok-4.6 在代理框架中的性能测试 · 跟踪 4 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Commentary
用户对模型性能的测试和评论,并非官方发布或基准测试。
Source corroboration
4 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
47 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [4]

  1. Mastodon — sigmoid.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @jun_song: 我在我的代理框架中测试了 DeepSeek-V4-Pro-0813 和 Grok-4.6。与基准测试(getes

    RT @jun_song: Ich habe DeepSeek-V4-Pro-0813 und Grok-4.6 in meiner Agent-Framework getestet. Die Leistung enttäuscht etwas im Vergleich zu den Benchmarks (getestet auf realen agentic Aufgaben, nicht auf Flappy Bird). Dennoch sind sie immer noch weit vor Opus-5, das derzeit aufgru…

  2. Mastodon — fosstodon.org TIER_1 Deutsch(DE) · [email protected] ·

    RT @thdxr: DeepSeek 在推理方面表现出色——它们的缓存命中率达到 96.56%,而我们第二好的提供商仅为 91.60%

    RT @thdxr: DeepSeek ist bei der Inferenz wahnsinnig gut – sie erreichen ein Cache-Verhältnis von 96,56 %, während unser zweitbester Anbieter nur 91,60 % schafft. Das mag nicht viel erscheinen, bedeutet aber, dass sie etwa die Hälfte der GPU-Zeit weniger benötigen. mehr auf Arint.…

  3. Mastodon — fosstodon.org TIER_1 Deutsch(DE) · [email protected] ·

    RT @jun_song: 我在我的代理框架上测试了 DeepSeek-V4-Pro-0813 和 Grok-4.6。与基准测试相比,性能有些令人失望(ge

    RT @jun_song: Ich habe DeepSeek-V4-Pro-0813 und Grok-4.6 auf meiner Agenten-Framework getestet. Die Leistung enttäuscht etwas im Vergleich zu den Benchmarks (getestet auf realen agentic Tasks, nicht auf Flappy Bird). Dennoch liegen sie immer noch weit vor Opus-5, das aufgrund von…

  4. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @jun_song: 我在我的代理框架中测试了 DeepSeek-V4-Pro-0813 和 Grok-4.6。与基准测试相比,性能有些令人失望 (getes

    RT @jun_song: Ich habe DeepSeek-V4-Pro-0813 und Grok-4.6 in meiner Agent-Framework getestet. Die Leistung enttäuscht etwas im Vergleich zu den Benchmarks (getestet auf realen agentic Tasks, nicht Flappy Bird). Dennoch sind sie immer noch weit vor Opus-5, das derzeit aufgrund von …