A recent evaluation by Velrim, a company selling data extraction APIs, revealed that AI models frequently invent information for fields not present in source documents. Across six tested systems, including models from Mistral AI, Gemini, and OpenAI, an average of 17% of absent fields were fabricated with invented values. Mistral AI's models showed a particularly high fabrication rate of approximately 40%. While Velrim's own API performed comparably to bare models like Gemini in terms of accuracy, its higher cost is justified by its ability to measure and report these fabrication rates, offering a published error metric. AI
IMPACT This study highlights a critical flaw in current LLM extraction capabilities, suggesting a need for improved hallucination mitigation and potentially impacting the reliability of AI-driven data processing.
RANK_REASON The cluster reports on a benchmark evaluation of AI model performance regarding data fabrication, which is a research finding. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →