PulseAugur
实时 07:34:17
English(EN) GPT-5.6 Sol beats Claude Fable 5 by 13.1 points on Agents' Last Exam — but there are some real caveats worth knowing

GPT 5.6-Sol 在代理任务上优于 Claude Fable-5,但仍存疑问

据报道,一款新模型 GPT 5.6-Sol 在长时序代理任务上超越了 Anthropic 的 Claude Fable-5,在 Agents' Last Exam 上的得分分别为 53.6 和 40.5。尽管在性能和价格上占有优势,但一些评论者认为 Fable-5 在架构推理和规划方面可能仍然表现出色。OpenAI 未发布 GPT 5.6-Sol 的长上下文回忆数字的决定也引发了疑问,特别是考虑到 Grok 和 Google 等其他主要实验室最近发布了自己的前沿模型。 AI

影响 为代理任务设定了新的基准,可能影响未来的模型开发和用户采用策略。

排序理由 前沿实验室模型发布及基准测试结果。[lever_c_demoted from frontier_release: ic=1 ai=1.0]

在 r/Anthropic 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

GPT 5.6-Sol 在代理任务上优于 Claude Fable-5,但仍存疑问

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
前沿实验室模型发布及基准测试结果。[lever_c_demoted from frontier_release: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
62 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [1]

  1. r/Anthropic TIER_1 English(EN) · /u/RajmaChawala ·

    GPT-5.6 Sol 在代理商的最后一次考试中以 13.1 分的优势击败 Claude Fable 5 — 但有一些真正的注意事项值得了解

    <!-- SC_OFF --><div class="md"><p>Yesterday's launch was genuinely interesting. Sol scores 53.6 vs Fable 5's 40.5 on long-horizon agentic tasks. On Terminal-Bench 2.1 it hits 88.8%. The pricing is also lower than Fable 5 for comparable or better agentic performance.</p> <p>BUT — …