PulseAugur
实时 22:43:38
English(EN) We analyzed GLM-5.3 and Claude Fable 5 on DeepSWE.

GLM-5.3 在编码任务上挑战 GPT-5.6 Sol 和 Claude Fable 5

Together AIDeepSWE 软件工程任务上对其 GLM-5.3 模型与 OpenAI 的 GPT-5.6 Sol 和 Anthropic 的 Claude Fable 5 进行了基准测试。GLM-5.3 表现出竞争力,在单次准确率上略微落后于 GPT-5.6 Sol,但在多次尝试和显著更低的成本下超越了它。与 Claude Fable 5 相比,GLM-5.3 在准确率上相当,但成本却低了五倍多,两种模型都表现出相似的失败模式和高度的任务一致性。 AI

影响 GLM-5.3 的表现表明,开放权重模型在软件工程等特定任务上正在缩小与前沿模型的差距,并提供了显著的成本优势。

排序理由 对多个 LLM 在特定任务上的比较基准测试。

在 X — Together (inference / OSS) 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

GLM-5.3 在编码任务上挑战 GPT-5.6 Sol 和 Claude Fable 5

报道来源 [3]

  1. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    我们分析了 DeepSWE 上的 GLM-5.3 和 Claude Fable 5。

    We analyzed GLM-5.3 and Claude Fable 5 on DeepSWE. GLM-5.3 MATCHED Fable 5 on single-shot solve rate at about one-fifth the cost, then pulled ahead with multiple attempts. The economics get interesting fast when retries are cheap 👇

  2. Together AI blog TIER_1 English(EN) ·

    GLM-5.3 对比 GPT-5.6 Sol 在 DeepSWE 上的表现:成本、编码和路由

    We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.

  3. Together AI blog TIER_1 English(EN) ·

    GLM-5.3 对比 Claude Fable 5 在 DeepSWE 上:成本、编码和路由

    We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.