A user benchmarked ten large language models on their ability to construct towers in a physics simulation, with Anthropic's Claude Opus 5 emerging as the winner. The benchmark involved placing blocks via a tool API, with noise introduced for precision in position or velocity. Claude Opus 5 achieved the highest standing height by strategically ending attempts to preserve tall structures, outperforming models like GPT-5.5 and DeepSeek V4 Flash. AI
IMPACT Demonstrates advanced reasoning and tool use capabilities in LLMs for complex physical simulations.
RANK_REASON User-generated benchmark of multiple LLMs on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
- Claude
- Claude Fable 5
- Claude Haiku 4.5
- Claude Opus 5
- Claude Sonnet 5
- DeepSeek V4 Flash
- GLM-5.2
- GPT-5.4 mini
- GPT-5.5
- GPT-5.6 Sol
- Kimi K3
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →