OmniDocBench
PulseAugur coverage of OmniDocBench — every cluster mentioning OmniDocBench across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
New DocPO framework enhances document parsing with tailored rewards
Researchers have introduced DocPO, a new framework for document parsing that utilizes reinforcement learning with tailored step-aware rewards. This approach aims to improve accuracy in complex document parsing tasks by …
-
LayoutLite module boosts OCR efficiency by reducing visual tokens
Researchers have developed LayoutLite, a novel module designed to enhance the efficiency of optical character recognition (OCR) systems that utilize vision-language models. This plug-and-play module operates by performi…
-
PaddlePaddle releases HPD-Parsing model with record throughput · 2 sources tracked
PaddlePaddle has released HPD-Parsing, a new lightweight document parsing model that utilizes a Hierarchical Parallel Decoding paradigm. This model achieves a new state-of-the-art score of 94.91% on the OmniDocBench v1.…
-
LLM judges unreliable for table recognition regeneration, study finds
A new research paper challenges the reliability of using Large Language Models (LLMs) as judges for evaluating and selecting outputs in closed-loop regeneration tasks, particularly in table recognition. The study found …
-
Youtu-Parsing model accelerates document analysis with novel decoding strategies
Researchers have introduced Youtu-Parsing, a novel document parsing model designed for efficient and high-performance content extraction. The system utilizes a Vision Transformer for feature extraction and a Youtu-LLM-2…
-
New OCR System Parses Long Financial Documents with Structure Awareness
Researchers have introduced LingDT-VL-OCR, a novel system designed for parsing ultra-long financial documents. This system aims to transform complex financial PDFs into accurate, structured outputs with auditable proven…
-
New framework enhances reading order inference for complex documents
Researchers have developed a novel training-free framework for inferring reading order in complex document layouts, particularly beneficial for digitizing historical manuscripts. This graph-based approach treats OCR tex…
-
Baidu releases Unlimited OCR, challenging long-context AI memory mechanisms · 1 source tracked
Baidu has open-sourced a new OCR model called Unlimited OCR, which excels at processing long documents by mimicking human reading habits. Unlike traditional OCR systems that process documents page by page and then stitc…
-
Open-source OCR models and benchmarks consolidated on Papers with Code
A new resource has been created to track open-source optical character recognition (OCR) models, consolidating information on top-performing models, benchmarks, and links to their papers and code. This initiative highli…
-
Baidu releases Unlimited OCR with constant KV cache for long documents
Baidu has released Unlimited OCR, a 3-billion-parameter Mixture-of-Experts model designed for efficient long-document parsing. The model utilizes Reference Sliding Window Attention (R-SWA) to maintain a constant KV cach…
-
ABot-OCR model transcribes pages directly to Markdown
Researchers have introduced ABot-OCR, a novel end-to-end vision-language model designed for direct transcription of page images into Markdown. This approach bypasses the need for complex modular systems by processing th…
-
New method enhances VLM document layout understanding
Researchers have developed a new method to improve how Vision-Language Models (VLMs) understand document layouts, particularly for documents with structures not seen during training. The approach pre-resolves layout inf…
-
New PureDocBench benchmark reveals document parsing is far from solved
Researchers have introduced PureDocBench, a new benchmark for document parsing that addresses issues with the existing OmniDocBench dataset, which suffers from annotation errors and potential contamination. PureDocBench…
-
RTPrune boosts DeepSeek-OCR inference speed by 1.23x with novel token pruning
Researchers have developed RTPrune, a novel two-stage token pruning method designed to enhance the efficiency of DeepSeek-OCR inference. This method mimics the model's two-stage reading process, first prioritizing high-…