DeepSeek has released its V4 Flash 0731 model, featuring three reasoning settings: Low, High, and Max. The model achieved impressive scores on the ARC Prize benchmark, a test for abstract reasoning in AI systems. Notably, in its Max setting, V4 Flash 0731 scored 89.0% on ARC-AGI-1 at a cost of $0.02 per task and 61.4% on the more challenging ARC-AGI-2 at $0.04 per task. This release highlights DeepSeek's continued focus on providing competitive performance at a fraction of the cost of Western models, with the model and its associated paper available on Hugging Face and arXiv. AI
IMPACT Sets new SOTA on abstract reasoning benchmarks, offering a cost-effective alternative to Western models.
RANK_REASON Frontier lab model release with benchmark results. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →