Mistral Large 4 has demonstrated a performance of 61.7% on the DeepSWE v1.1 benchmark for coding tasks. This score places it behind both Chinese open-weight models, which achieved approximately 69%, and leading closed-source models from OpenAI, Google, and Anthropic, which scored around 74%. The data indicates that performance gaps continue to exist, even as open-weight models see increased development and availability. AI
IMPACT Mistral Large 4's performance on coding benchmarks indicates areas for improvement compared to leading closed-source and some open-weight models.
RANK_REASON The item reports on benchmark results for a specific AI model, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →