PulseAugur
中
实时 07:35:07
English(EN) TWIST : A benchmark where the model can only see the Rubik's cube through screenshots. Opus 5 solved it - 44 minutes, 99% of that was thinking.

Anthropic 的 Claude Opus 5 通过截图基准测试解决了魔方

Anthropic 的 Claude Opus 5 使用一个名为 TWIST 的新基准测试成功解决了复杂的魔方谜题。该基准测试要求 AI 解释魔方的截图并发出指令,模拟真实场景,而无需直接访问魔方的状态。Claude Opus 5 在 44 分钟内完成了 20 步打乱,使用了 73 张截图和约 240,000 个 token,大部分时间花在“思考”上,而不是魔方操作。 AI

影响 展示了 LLM 在视觉推理和解决复杂问题方面的先进能力,可能对机器人技术和复杂任务自动化产生影响。

排序理由 该条目描述了一个新的基准测试和模型在该测试上的表现,属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/Anthropic 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Anthropic 的 Claude Opus 5 通过截图基准测试解决了魔方

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该条目描述了一个新的基准测试和模型在该测试上的表现,属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
70 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/Anthropic TIER_1 English(EN) · /u/Various-Affect4841 ·

    TWIST:一个模型只能通过截图看到魔方块的基准测试。Opus 5 解决了它——44分钟,其中99%是思考时间。

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1v6hcgm/twist_a_benchmark_where_the_model_can_only_see/"> <img alt="TWIST : A benchmark where the model can only see the Rubik's cube through screenshots. Opus 5 solved it - 44 minutes, 99% of that was thinking…