NaviDC-OCR is a new vision-language framework designed to improve document parsing by addressing challenges in handling distorted camera-captured documents and enhancing structural reasoning. The framework incorporates deformation-aware learning, adaptive layout sampling, and a content-structure decoupled training strategy. Experiments show NaviDC-OCR achieves state-of-the-art results on several benchmarks, including OmniDocBench v1.6 and Wild-OmniDocBench, and secured first place in the ICDAR 2026 Sci-ImageMiner Challenge. AI
IMPACT This framework could improve the accuracy and structural understanding of documents processed by AI systems.
RANK_REASON The item describes a new research paper detailing a novel framework for document parsing. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
- ICDAR 2026 Sci-ImageMiner Challenge
- NaviDC-OCR
- OmniDocBench v1.6
- PureDocBench
- vision-language model
- Wild-OmniDocBench
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →