PulseAugur
EN
LIVE 15:34:37

AI models fabricate 17% of absent data fields, study finds

A recent evaluation by Velrim, a company selling data extraction APIs, revealed that AI models frequently invent information for fields not present in source documents. Across six tested systems, including models from Mistral AI, Gemini, and OpenAI, an average of 17% of absent fields were fabricated with invented values. Mistral AI's models showed a particularly high fabrication rate of approximately 40%. While Velrim's own API performed comparably to bare models like Gemini in terms of accuracy, its higher cost is justified by its ability to measure and report these fabrication rates, offering a published error metric. AI

IMPACT This study highlights a critical flaw in current LLM extraction capabilities, suggesting a need for improved hallucination mitigation and potentially impacting the reliability of AI-driven data processing.

RANK_REASON The cluster reports on a benchmark evaluation of AI model performance regarding data fabrication, which is a research finding. [lever_c_demoted from research: ic=1 ai=1.0]

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

AI models fabricate 17% of absent data fields, study finds

How we ranked this

Signal score
54 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The cluster reports on a benchmark evaluation of AI model performance regarding data fabrication, which is a research finding. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Velrim ·

    Extraction APIs invented values for 17% of the fields that aren't in the document. Ours included.

    <p>Velrim ran this comparison, and we sell one of the six APIs in the table. Every raw output, request ID, and the scoring CLI are in the repo (<a href="https://github.com/velrimhq/velrim-eval" rel="noopener noreferrer">https://github.com/velrimhq/velrim-eval</a>).</p> <h2> Why w…