PulseAugur
实时 13:35:25
English(EN) We compared how far the same budget goes with DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE.

DeepSeek-V4 Flash 以成本效益挑战 GPT-5.6 Luna 编码基准

Together AI 发布了对 DeepSeek-V4 FlashGPT-5.6 Luna 在 DeepSWE 编码基准上的对比分析。虽然 GPT-5.6 Luna 在所有质量指标上都表现出卓越的性能,但 DeepSeek-V4 Flash 被证明具有显著的成本效益。分析表明,采用级联方法,优先使用 DeepSeek-V4 Flash,仅在必要时升级到 GPT-5.6 Luna,可以以比单独使用 GPT-5.6 Luna 更低的成本实现更高的准确性。 AI

影响 通过结合更便宜、功能强大的模型和更昂贵、更强大的模型,提出了在编码任务中利用 LLM 的成本效益策略。

排序理由 对两个模型在特定基准上的对比分析。

在 X — Together (inference / OSS) 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

DeepSeek-V4 Flash 以成本效益挑战 GPT-5.6 Luna 编码基准

报道来源 [3]

  1. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    我们将预算花在 DeepSeek V4 Flash 和 GPT-5.6 Luna 在 DeepSWE 上的效果进行了对比。

    We compared how far the same budget goes with DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. Two DeepSeek V4 Flash attempts solved MORE tasks than one Luna attempt at roughly one-third the cost. https://t.co/yzLS3E3v9E

  2. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    我们分析了 DeepSWE 上的 DeepSeek V4 Flash 和 GPT-5.6 Luna。

    We analyzed DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. A DeepSeek-first cascade with test-suite verification solved MORE tasks than Luna alone at 37% lower cost per task. https://t.co/fMUH9ulrVR

  3. Together AI blog TIER_1 Nederlands(NL) ·

    DeepSeek-V4 Flash 0731 对比 GPT-5.6 Luna 在 DeepSWE 上的成本与编码表现

    We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.