Researchers have introduced ExtractBench, a new benchmark designed to evaluate schema-guided document extraction capabilities in enterprise settings. This benchmark assesses agents on value accuracy, record completeness, grounding, and cost, utilizing a dataset of 4,869 pages across 370 enterprise documents. Initial results show that while commercial VLMs struggle with long documents, coding agents are accurate but expensive. LlamaExtract Agentic Plus emerged as the top performer, offering comparable accuracy to coding agents at a significantly lower cost. AI
IMPACT This benchmark could drive improvements in enterprise AI agents for accurate and cost-effective data extraction.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for AI capabilities.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →