CursorBench 4.0 results show Opus 5.5 Max achieving the highest score at 57.8%, though at a cost of $13.43 per task. Sonnet 5.5 Max and Haiku 5.5 Max offered better cost-efficiency, scoring 55.5% ($7.05) and 48.4% ($1.12) respectively. The performance differences between the models may not be statistically significant. AI
IMPACT Provides comparative performance and cost data for AI models, aiding in selection for specific applications.
RANK_REASON Benchmark results for AI models. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →