PulseAugur
EN
LIVE 10:46:36

NaviDC-OCR framework enhances document parsing with deformation-aware learning

Researchers have introduced NaviDC-OCR, a novel framework designed to enhance document parsing across both digital and camera-captured documents. This system addresses limitations in existing methods by incorporating deformation-aware learning to better handle geometric distortions and employing an adaptive sampling mechanism for complex layouts. NaviDC-OCR also utilizes a content-structure decoupled learning strategy to explicitly model formulas and tables, leading to improved structured representation. The framework has demonstrated state-of-the-art performance on several benchmarks, including OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, and secured first place in the ICDAR 2026 Sci-ImageMiner Challenge. AI

IMPACT This framework could improve the accuracy and efficiency of extracting information from diverse document types, benefiting applications that rely on structured data.

RANK_REASON The item is a research paper detailing a new framework for document parsing. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

NaviDC-OCR framework enhances document parsing with deformation-aware learning

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Peng Cai, Zhaofan Zou, Shifa Liu, Yikun Wang, Jiawei Tang, Kaicheng Yang, Meng Tong, Zhongjiang He, Hao Sun ·

    NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

    arXiv:2608.12898v1 Announce Type: cross Abstract: Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing appro…