PulseAugur
实时 08:16:24
English(EN) TWIST : A benchmark where the model can only see the Rubik's cube through screenshots. Opus 5 solved it - 44 minutes, 99% of that was thinking.

Anthropic 的 Claude Opus 5 通过截图基准测试解决了魔方

AnthropicClaude Opus 5 使用一个名为 TWIST 的新基准测试成功解决了复杂的魔方谜题。该基准测试要求 AI 解释魔方的截图并发出指令,模拟真实场景,而无需直接访问魔方的状态。Claude Opus 5 在 44 分钟内完成了 20 步打乱,使用了 73 张截图和约 240,000 个 token,大部分时间花在“思考”上,而不是魔方操作。 AI

影响 展示了 LLM 在视觉推理和解决复杂问题方面的先进能力,可能对机器人技术和复杂任务自动化产生影响。

排序理由 该条目描述了一个新的基准测试和模型在该测试上的表现,属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/Anthropic 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Anthropic 的 Claude Opus 5 通过截图基准测试解决了魔方

报道来源 [1]

  1. r/Anthropic TIER_1 English(EN) · /u/Various-Affect4841 ·

    TWIST:一个模型只能通过截图看到魔方块的基准测试。Opus 5 解决了它——44分钟,其中99%是思考时间。

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1v6hcgm/twist_a_benchmark_where_the_model_can_only_see/"> <img alt="TWIST : A benchmark where the model can only see the Rubik's cube through screenshots. Opus 5 solved it - 44 minutes, 99% of that was thinking…