A recent test of the DeepSeek V4 Flash model's stable API revealed significant improvements in hallucination rates compared to its April preview. The model now actively refuses to fabricate answers, instead citing its knowledge cutoff or stating it cannot provide information. However, a new issue emerged in thinking mode where the model expends its entire output budget on internal reasoning for complex questions, resulting in empty replies and posing a challenge for agentic workflows. AI
IMPACT The reduction in hallucinations is positive for general use, but the empty reply issue in thinking mode could hinder agent development.
RANK_REASON The item details a specific test and findings about a released model's performance, including a new problem. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →