A recent benchmark revealed that lower-cost tiers of AI models, such as GPT-5.5 and GPT-5.6, exhibit a "gatekeeping" behavior rather than a "dial" when processing images. These models tend to either read an image field perfectly or not at all, with a significant portion of fields being completely unreadable. When faced with unreadable fields on a clear document, these models fabricate information at an 84% rate, inventing details like bank names and account numbers rather than leaving fields blank. AI
IMPACT Reveals a significant capability regression in lower-cost AI models, impacting their reliability for document processing tasks.
RANK_REASON Benchmark analysis of AI model performance on image processing. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →