PulseAugur
实时 01:46:58
English(EN) In our DeepSWE analysis, Kimi K3 Max came close to Fable 5 xhigh on Pass@1 while costing about one-third as much per rollout.

Kimi K3 Max 在编码任务上以更低的成本媲美 Fable 5 xhigh

Together 的 DeepSWE 分析显示,Kimi K3 Max 在 Pass@1 指标上的表现与 Fable 5 xhigh 相当。值得注意的是,Kimi K3 Max 的成本效益显著更高,每次部署的成本约为三分之一,每美元解决的任务数量是其 2.8 倍。这凸显了成本效益在大型模型部署中的重要性。 AI

影响 强调了大规模运营中人工智能模型部署的成本效益。

排序理由 该条目详细介绍了人工智能模型在特定任务(DeepSWE 分析)上的基准比较,属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]

在 X — Together (inference / OSS) 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Kimi K3 Max 在编码任务上以更低的成本媲美 Fable 5 xhigh

报道来源 [1]

  1. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    在我们的DeepSWE分析中,Kimi K3 Max在Pass@1上的表现接近Fable 5 xhigh,而每次推出的成本仅为其三分之一。

    In our DeepSWE analysis, Kimi K3 Max came close to Fable 5 xhigh on Pass@1 while costing about one-third as much per rollout. That resulted in 2.8× more solved tasks per dollar. For teams running models at scale, the cost per successful task is often the more useful comparison.…