PulseAugur
EN
LIVE 19:01:40

Real-world email test reveals LLM formatting failures in cheaper models

A company that uses AI to generate cold emails found that cheaper models like DeepSeek V4 Flash, Gemini Flash Lite, and GLM failed to maintain proper email formatting, specifically collapsing paragraphs into a single block. This issue was not detected by standard benchmarks or the structured output schema, leading to unsendable emails. The company's default model, GPT-5 Mini, performed best in this real-world test, highlighting the importance of robust structured output capabilities beyond basic accuracy scores. AI

IMPACT Highlights the need for LLM evaluations that go beyond benchmarks to assess real-world usability and formatting capabilities.

RANK_REASON This is a commentary on the practical performance of LLMs in a specific application, rather than a release or research paper.

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Real-world email test reveals LLM formatting failures in cheaper models

How we ranked this

Signal score
11 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Commentary
This is a commentary on the practical performance of LLMs in a specific application, rather than a release or research paper.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Julian M. Wagner ·

    We replayed real cold email prompts through 7 LLMs. DeepSeek, Gemini Lite and GLM failed in a way no benchmark shows

    <p>Every email our platform sends is written by a model. Not filled in from a template: written, per recipient, from the recipient's website and a customer's brief. That makes model choice a cost line, not a taste question. A model that costs half as much per token saves real mon…