PulseAugur
实时 04:12:58
(MK) GPT-5.6 Sol поставила рекорд Terminal-Bench: разбираем, чем она кодит

OpenAI 的 GPT-5.6 Sol 代码模型因基准测试声明面临审查

OpenAI 于 2026 年 7 月 9 日发布了 GPT-5.6 Sol,这是一款专注于编码和复杂推理的新旗舰 AI 模型。虽然 OpenAI 声称在 Terminal-Bench 2.1 上取得了 88.8% 和 91.9% 的惊人基准分数,但这些结果来自内部测试,尚未在独立排行榜上反映出来。独立评估显示,在 SWE-Bench Pro 等基准测试中,Sol 落后于 Claude Fable 5 等竞争对手,尽管另一项 METR 评估指出 Sol 的测试通过率创下了纪录。 AI

影响 在编码基准测试上设定了新的 SOTA;给 Anthropic 带来压力,要求其做出回应。

排序理由 前沿实验室模型发布,附带系统卡。[lever_c 从 frontier_release 降级:ic=1 ai=1.0]

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

OpenAI 的 GPT-5.6 Sol 代码模型因基准测试声明面临审查

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
前沿实验室模型发布,附带系统卡。[lever_c 从 frontier_release 降级:ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
45 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [1]

  1. dev.to — LLM tag TIER_1 (MK) · Promptra Team ·

    GPT-5.6 Sol 在 Terminal-Bench 上创下新纪录:我们来分析一下它都写了些什么代码

    <p><em>Применить: вечер на оценку модели под свой стек · Уровень: средний · Чтение: ~27 минут · Данные проверены на 10 июля 2026</em></p> <blockquote> <p><strong>Главное.</strong> GPT-5.6 Sol - флагманская нейросеть для кода из семейства OpenAI, вышедшего публично 9 июля 2026. Пр…