PulseAugur
EN
LIVE 07:27:57

Anthropic's Claude Opus 5 solves Rubik's Cube via screenshot benchmark

Anthropic's Claude Opus 5 has successfully solved a complex Rubik's Cube puzzle using a new benchmark called TWIST. This benchmark requires the AI to interpret screenshots of the cube and issue commands, simulating a real-world scenario without direct access to the cube's state. Claude Opus 5 completed the 20-move scramble in 44 minutes, utilizing 73 screenshots and approximately 240,000 tokens, with the majority of the time spent in 'thinking' rather than cube manipulation. AI

IMPACT Demonstrates advanced visual reasoning and problem-solving capabilities in LLMs, potentially impacting robotics and complex task automation.

RANK_REASON The item describes a new benchmark and a model's performance on it, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/Anthropic →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Anthropic's Claude Opus 5 solves Rubik's Cube via screenshot benchmark

COVERAGE [1]

  1. r/Anthropic TIER_1 English(EN) · /u/Various-Affect4841 ·

    TWIST : A benchmark where the model can only see the Rubik's cube through screenshots. Opus 5 solved it - 44 minutes, 99% of that was thinking.

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1v6hcgm/twist_a_benchmark_where_the_model_can_only_see/"> <img alt="TWIST : A benchmark where the model can only see the Rubik's cube through screenshots. Opus 5 solved it - 44 minutes, 99% of that was thinking…