Alibaba's Qwen3.8-Max-0902 model has achieved the top position in the Code Arena's WebDev benchmark, scoring 1,691 points. This model also leads the Pareto frontier for cost-effectiveness, priced at $5 per million tokens. This performance places it above other models like HY4 Preview and significantly undercuts the cost of Claude Opus, making advanced agent loops more accessible for developers. AI
IMPACT Sets a new standard for cost-effective performance in coding benchmarks, potentially lowering the barrier for advanced AI agent deployment.
RANK_REASON Model benchmark achievement and cost-effectiveness analysis.
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →