A recent experiment involving 360 rounds of AI novel continuation revealed three primary failure modes across nine different models. These failures include models restarting the narrative from the beginning, getting stuck in repetitive loops of their own generated text, or rewinding the plot to an earlier point. The study found that even advanced models like Gemini-3.1 Pro and GPT-5.6 Terra exhibited these issues, highlighting a disconnect between surface-level prose mimicry and long-range narrative coherence. Models like DeepSeek V4 Flash and V4 Pro showed stronger performance, with the former excelling in a fantasy epic and the latter in a palace-intrigue novel, suggesting that model performance is genre-dependent. AI
IMPACT Highlights limitations in current LLMs for long-form creative writing, suggesting a need for better context management and coherence in future models.
RANK_REASON The item details a research experiment and its findings on AI model capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- DeepSeek V4 Flash
- DeepSeek V4 Pro
- Gemini-3.1 Pro
- GLM 5.2
- GPT-5.6 Terra
- Grok 4.5
- Kimi K2.6
- The Legend of Zhen Huan
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →