A new study investigates the effectiveness of Vision-Language Models (VLMs) for Optical Character Recognition (OCR) on Arabic manuscripts. Researchers found that no single approach consistently outperforms others across various datasets, including historical manuscripts, aged print, and handwriting. An OCR-conditioned VLM correction method shows promise, particularly when the initial OCR output is recoverable and visually similar to the original text. However, this method can degrade performance if the OCR prior is mismatched or misleading, such as with certain manuscript scripts or handwriting styles. The findings suggest an adaptive workflow that routes pages based on script, OCR recoverability, and failure indicators. AI
IMPACT Findings suggest an adaptive OCR-VLM workflow is needed, potentially improving accuracy for specific Arabic manuscript types.
RANK_REASON The cluster contains two academic papers presenting research findings on OCR and VLMs.
- Arabic manuscript OCR
- Hugging Face
- KHATT
- Maghribi manuscripts
- Muharaf
- Naskh manuscripts
- OCR-conditioned VLM correction
- SaudiHeritage-OCR
- Tesseract
- Vision--Language Models
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →