optical character recognition
PulseAugur coverage of optical character recognition — every cluster mentioning optical character recognition across labs, papers, and developer communities, ranked by signal.
12 day(s) with sentiment data
-
New method deciphers how VLMs verbalize image semantics using OCR heads
Researchers have developed a method to understand how Vision-Language Models (VLMs) process image semantics, focusing specifically on their optical character recognition (OCR) capabilities. By identifying specific atten…
-
Quantum-Inspired Transformer (QiT) Advances Visual Recognition
Researchers have developed QiT, a Quantum-inspired Transformer model for visual recognition tasks. QiT leverages structural ideas from quantum models, such as angle-inspired encoding and periodic feature self-attention,…
-
LandingAI launches ADE Gen2 with atomic grounding and agent-ready JSON
LandingAI has launched the second generation of its Agentic Document Extraction (ADE Gen2) system, built upon its new DPT-3 model family. This update focuses on improving output structure, grounding, and cost-effectiven…
-
AI workflow screens South African foods for sodium compliance
Researchers have developed an OCR-enabled workflow to screen South African packaged foods for sodium content against regulatory limits. The system combines region detection, optical character recognition, and vision-lan…
-
Vision-language models evaluated for document extraction, revealing trade-offs
A new research paper explores the trade-offs involved in using vision-language models (VLMs) for extracting structured data from business documents. The study evaluated eleven systems, including commercial offerings lik…
-
VideoXAgent tackles long video understanding with online agent harness
Researchers have developed VideoXAgent, an online harness designed for understanding long videos. This system plans tasks, uses specialized tools like VLMs, OCR, and ASR on demand, and aggregates evidence to answer quer…
-
MinerU leads Document AI & OCR rankings on Olud Pulse
MinerU has achieved the top ranking on the Olud Pulse for Document AI and OCR, scoring 89 out of 100. This score represents a slight decrease from its previous performance. The Olud Pulse provides a comparative analysis…
-
OCR enhanced by layout models and table detection for business documents
Optical character recognition (OCR) can extract text from documents, but it struggles with complex layouts and ensuring data accuracy. Advanced techniques like layout models and table detection are crucial for transform…
-
Image annotation costs in 2026 to vary by complexity and type
The cost of image annotation in 2026 will depend on several factors, including the complexity of the task, volume of data, quality assurance processes, required expertise, and turnaround time. Specific annotation types …
-
Document image classification explanations improved by domain-aware segmentation
A new study published on arXiv investigates the impact of segmentation choices on the reliability of LIME explanations for document image classification. Researchers found that standard superpixel-based segmentations, c…
-
New MDPO method enhances LLM-based historical entity linking
Researchers have developed a new method called Multi-Negative Direct Preference Optimization (MDPO) to improve historical entity linking using large language models. Unlike previous approaches that only considered one n…
-
New AI agent accurately converts blueprints to simulation models
Researchers have developed BlueprintAgent (BPA), a novel multimodal agent designed to accurately convert scanned structural blueprints into simulation-ready models. Unlike previous methods that relied on direct promptin…
-
FlowCPO introduces unified divergence view for preference alignment in flow models
Researchers have introduced FlowCPO, a novel method for aligning flow and diffusion models using an offline forward-KL objective. This approach unifies existing online reinforcement learning and offline preference optim…
-
LLMs advance invoice extraction but don't automate accounts payable
While Large Language Models (LLMs) have significantly improved invoice extraction by enabling template-free processing and handling document messiness, they do not fully automate accounts payable. LLMs excel at the init…
-
New LeakageBench benchmark reveals persistent PII risks in document redaction
A new benchmark called LeakageBench has been developed to assess the risk of personally identifiable information (PII) leakage from document images. The benchmark, which includes 500 document images with over 11,000 GDP…
-
New benchmark VeriOCRBench tests MLLMs for OCR task verification
Researchers have introduced VeriOCRBench, a new benchmark designed to evaluate the task verification capabilities of Multimodal Large Language Models (MLLMs) in optical character recognition (OCR) scenarios. This benchm…
-
OCR fine-tuning unlocks archaeological pottery metadata extraction
Researchers have developed a method to extract metadata from archaeological pottery records using OCR technology, addressing the challenge of manually transcribing handwritten documents. A new dataset, CENTURIA, contain…
-
New ClearText-Video dataset probes MLLM text-reading in low-quality videos
Researchers have introduced ClearText-Video (CTVid), a new large-scale dataset designed to evaluate how multimodal large language models (MLLMs) handle text-centric video understanding under varying quality conditions. …
-
New benchmark tests MLLMs' meta-reasoning in text-rich images
Researchers have introduced the OCR-MetaReasoning Benchmark, a new evaluation tool designed to assess the meta-reasoning capabilities of multimodal large language models (MLLMs) in understanding images containing text. …
-
New framework boosts OCR for low-resource languages
Researchers have developed a new framework called PSMC to improve Optical Character Recognition (OCR) for low-resource languages. Traditional fine-tuning methods struggle with limited data, but PSMC leverages a cross-sc…