DeepSeek V4 Pro 0813 has achieved notable scores on several benchmarks, including 92.8% on GPQA and 80.3% on Long Context Reasoning. The model also reached 51% on SciCode and 41% on Humanity's Last Exam. These results were independently measured, indicating a performance of 81.1 tokens per second and 18.3 intelligence points per dollar. AI
IMPACT Sets new benchmark scores for reasoning and long context, potentially influencing future model development.
RANK_REASON The item reports benchmark scores for an AI model, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →