A new benchmark, Terminal-Bench 4.0, highlights the significant cost differences between top-performing AI models, even when their performance scores are nearly identical. GPT-6 Astra, running through Codex, achieved a score of 58.2% at a cost of approximately $3,300 for a full run, while Claude Fable 5.1, using Claude Code, scored 57.9% but cost nearly double at $6,200. The analysis suggests that developers should prioritize cost-effectiveness alongside performance, as minor score differences may not justify substantial price increases, and models like Gemini 3.8 Flash offer strong value for their cost. AI
IMPACT Highlights the critical need for cost-performance analysis in LLM selection, suggesting minor performance gains may not justify significant price increases.
RANK_REASON Analysis of a new benchmark with detailed cost-performance data. [lever_c_demoted from research: ic=1 ai=1.0]
- Claude Code
- Claude Fable 5.1
- codex
- Gemini 3.8 Flash
- GLM-5
- GPT-6 Astra
- Hermes Agent
- Nokka
- OpenAI
- Sonnet 5
- Tokencost
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →