A company that uses AI to generate cold emails found that cheaper models like DeepSeek V4 Flash, Gemini Flash Lite, and GLM failed to maintain proper email formatting, specifically collapsing paragraphs into a single block. This issue was not detected by standard benchmarks or the structured output schema, leading to unsendable emails. The company's default model, GPT-5 Mini, performed best in this real-world test, highlighting the importance of robust structured output capabilities beyond basic accuracy scores. AI
IMPACT Highlights the need for LLM evaluations that go beyond benchmarks to assess real-world usability and formatting capabilities.
RANK_REASON This is a commentary on the practical performance of LLMs in a specific application, rather than a release or research paper.
- AI SDK
- DeepSeek V4 Flash
- Gemini 3 Flash
- Gemini Flash Lite
- General Language Model
- GPT-5 Mini
- GPT-5 Nano
- Kimi K2.5
- Zod
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →