A new research paper explores the trade-offs involved in using vision-language models (VLMs) for extracting structured data from business documents. The study evaluated eleven systems, including commercial offerings like GPT-5 and open-source models such as Claude Sonnet 4.5, on a dataset of synthetic checks. Fine-tuning open-source VLMs significantly improved their performance, surpassing commercial systems in some cases, while GPT-5 led in overall accuracy and Claude Sonnet 4.5 struggled with date extraction. The research also introduces a framework to help practitioners select the most suitable approach based on factors like quality, latency, governance, and volume. AI
IMPACT Provides practical guidance for selecting VLMs in document extraction, highlighting performance and cost trade-offs.
RANK_REASON The cluster contains a research paper detailing an evaluation of vision-language models for a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Claude Sonnet 4.5
- DagsHub
- Gotit.pub
- GPT-5
- Hugging Face
- optical character recognition
- regular expression
- ScienceCast
- vision-language model
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →