A comparison of AI models playing the game Liar's Dice revealed that Claude Code, specifically Claude Opus 5, outperformed Codex CLI (gpt-5.6-sol). In three best-of-three series, Claude Code won each match 2-0, demonstrating superior challenge call accuracy. The setup ensured fair play by using a dedicated engine and MCP servers, preventing models from seeing each other's dice or using side channels. AI
IMPACT Demonstrates differences in strategic reasoning and bluffing capabilities between AI models in a game context.
RANK_REASON Comparison of AI model performance on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →