PulseAugur
实时 11:49:53
English(EN) GLM-5.3 Arrives Hours After Tang Jie's 'sooooooon' — Zhipu's Coding and Security Model Doubles SWE-Marathon, Tops CyberGym

Zhipu AI 发布 GLM-5.3,在编码和安全基准测试中表现出色

Zhipu AI 发布了其新模型 GLM-5.3,该模型专为编码和安全任务设计。该模型在基准测试中的性能显著提高,在 SWE-Marathon 上的得分翻了一番多,在 Terminal Bench 3.0 上的得分翻了五倍。GLM-5.3 在 CyberGym 基准测试中也取得了最高分。 AI

影响 在编码和安全基准测试中设定了新的 SOTA(State-of-the-Art),可能影响未来专业模型的发展。

排序理由 Frontier-lab 模型发布,附带系统卡。[lever_c_demoted from frontier_release: ic=1 ai=1.0]

在 Pandaily 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Zhipu AI 发布 GLM-5.3,在编码和安全基准测试中表现出色

报道来源 [1]

  1. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    GLM-5.3 在唐杰“sooooooon”数小时后发布 — Zhipu 的编码和安全模型将 SWE-Marathon 成绩翻倍,在 CyberGym 中名列前茅

    Three days after a user teased Zhipu AI chief scientist Tang Jie on X about GLM-5.3, he replied 'sooooooon', and hours later the model was live. Focused solely on coding and security, GLM-5.3 more than doubled SWE-Marathon to 42.5, quintupled Terminal Bench 3.0 to 28.3, and score…