PulseAugur
实时 02:39:07
English(EN) OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness

OpenAI 的 GPT-5.6 Sol 基准测试声明因自定义测试环境而受到质疑

OpenAI 声称其新的 GPT-5.6 Sol 模型在 ARC-AGI-3 基准测试上可以超越 Anthropic 的 Opus 5。然而,这一 38.3% 的优异分数是通过使用 OpenAI 的专有 API 功能实现的,包括保留推理和上下文压缩。在官方、提供商中立的 ARC-AGI-3 环境中进行测试时,GPT-5.6 Sol 的得分仅为 7.8%,而 Opus 5 在没有此类专门设置的情况下达到了 30.2%。 AI

影响 凸显了在大型语言模型公平和标准化基准测试方面持续存在的挑战。

排序理由 该集群报道了 OpenAI 关于模型在基准测试上表现的声明,但重点是围绕测试方法论的争议,而不是官方发布或研究论文。

在 The Decoder 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

OpenAI 的 GPT-5.6 Sol 基准测试声明因自定义测试环境而受到质疑

报道来源 [4]

  1. The Decoder TIER_1 English(EN) · Matthias Bastian ·

    OpenAI 声称 GPT-5.6 Sol 在 ARC-AGI-3 上击败 Opus 5,凭借其最新 API 和两个额外设置

    <p><img alt="" class="attachment-full size-full wp-post-image" height="768" src="https://the-decoder.com/wp-content/uploads/2026/06/openai_gpt56_sol.png" style="height: auto; margin-bottom: 10px;" width="1376" /></p> <p> OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol s…

  2. The Decoder TIER_1 English(EN) · Matthias Bastian ·

    OpenAI声称GPT-5.6 Sol在ARC-AGI-3上击败Opus 5,但仅在其自有的定制测试框架下

    <p><img alt="" class="attachment-full size-full wp-post-image" height="768" src="https://the-decoder.com/wp-content/uploads/2026/06/openai_gpt56_sol.png" style="height: auto; margin-bottom: 10px;" width="1376" /></p> <p> OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol s…

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    OpenAI 表示 GPT-5.6 Sol 在 ARC-AGI-3 上击败 Opus 5,但仅限于使用其自有 API 功能。若无这些功能,得分降至 7.8%。公平基准测试仍是一项挑战

    OpenAI says GPT-5.6 Sol beats Opus 5 on ARC-AGI-3, but only with its own API features. Without them, the score drops to 7.8%. Fair benchmarking remains a challenge. Source: The Decoder AI https:// the-decoder.com/openai-claims- gpt-5-6-sol-beats-opus-5-on-arc-agi-3-with-its-lates…

  4. Mastodon — mastodon.social TIER_1 English(EN) · sipirtu ·

    OpenAI 表示 GPT-5.6 Sol 在 ARC-AGI-3 上通过自定义工具超越 Opus 5,但在官方测试中得分 7.8%。Opus 5 在没有特殊优化的情况下达到了 30.2%。

    OpenAI says GPT-5.6 Sol tops Opus 5 on ARC-AGI-3 with a custom harness but scores 7.8 percent in the official test. Opus 5 reached 30.2 percent without special setups. Source: The Decoder AI https:// the-decoder.com/openai-claims- gpt-5-6-sol-beats-opus-5-on-arc-agi-3-but-only-wi…