A new study published on arXiv explores the use of lightweight, open-source vision-language models (VLMs) for document OCR and structured JSON extraction. The research compares eight VLMs with up to 7 billion parameters, evaluating their performance in zero-shot, few-shot, and fine-tuning settings. The findings suggest that these smaller VLMs can offer a sustainable, private, and high-performing alternative to manual transcription or commercial systems, providing guidance for heritage institutions. AI
IMPACT Provides guidance for heritage institutions on using controlled, efficient VLMs for document digitization and data extraction.
RANK_REASON Research paper published on arXiv detailing a comparative study of VLMs for document processing. [lever_c_demoted from research: ic=1 ai=1.0]
- Approximate Normalized Levenshtein Similarity
- arXiv
- Character Error Rate
- JSON
- mAP-F1
- mean Average Precision F1
- optical character recognition
- vision-language model
- Vision--Language Models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →