Researchers have introduced NaviDC-OCR, a novel framework designed to enhance document parsing across both digital and camera-captured documents. This system addresses limitations in existing methods by incorporating deformation-aware learning to better handle geometric distortions and employing an adaptive sampling mechanism for complex layouts. NaviDC-OCR also utilizes a content-structure decoupled learning strategy to explicitly model formulas and tables, leading to improved structured representation. The framework has demonstrated state-of-the-art performance on several benchmarks, including OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, and secured first place in the ICDAR 2026 Sci-ImageMiner Challenge. AI
IMPACT This framework could improve the accuracy and efficiency of extracting information from diverse document types, benefiting applications that rely on structured data.
RANK_REASON The item is a research paper detailing a new framework for document parsing. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Hugging Face
- ICDAR 2026 Sci-ImageMiner Challenge
- NaviDC-OCR
- OmniDocBench v1.6
- PureDocBench
- Wild-OmniDocBench
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →