PulseAugur
EN
LIVE 16:44:17

NaviDC-OCR framework enhances document parsing with deformation-aware learning

NaviDC-OCR is a new vision-language framework designed to improve document parsing by addressing challenges in handling distorted camera-captured documents and enhancing structural reasoning. The framework incorporates deformation-aware learning, adaptive layout sampling, and a content-structure decoupled training strategy. Experiments show NaviDC-OCR achieves state-of-the-art results on several benchmarks, including OmniDocBench v1.6 and Wild-OmniDocBench, and secured first place in the ICDAR 2026 Sci-ImageMiner Challenge. AI

IMPACT This framework could improve the accuracy and structural understanding of documents processed by AI systems.

RANK_REASON The item describes a new research paper detailing a novel framework for document parsing. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

NaviDC-OCR framework enhances document parsing with deformation-aware learning

COVERAGE [1]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

    NaviDC-OCR is a unified vision-language framework that integrates deformation-aware learning, adaptive layout sampling, and decoupled content-structure training to improve document parsing accuracy and structural reasoning.