PulseAugur
EN
LIVE 07:28:52

EVICALC system achieves 8.61% accuracy in DocSem task, highlights OCR challenges

A system named EVICALC, designed for the DocSem shared task, achieved 8.61% joint accuracy on 1,730 tasks. The system processes PDFs by selecting passages, using a language model to generate arithmetic expressions, and then evaluating these expressions locally. A separate public-validation run yielded higher scores of 92.17% answer accuracy and 1.00 evidence F1, though these metrics are not directly comparable due to differing configurations. Analysis revealed that issues such as optical character recognition errors merging text blocks can lead to the system answering from unrelated content. AI

IMPACT Highlights challenges in integrating OCR and LLMs for complex document understanding tasks.

RANK_REASON The cluster contains a research paper detailing a system's performance and failure analysis on a specific task.

Read on arXiv cs.IR (Information Retrieval) →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

EVICALC system achieves 8.61% accuracy in DocSem task, highlights OCR challenges

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
The cluster contains a research paper detailing a system's performance and failure analysis on a specific task.
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
6 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [2]

  1. arXiv cs.CL TIER_1 English(EN) · Divya Godara, Sachin Gupta ·

    Evidence First, Arithmetic Second: A System Report and Failure Analysis for DocSem

    arXiv:2609.39013v1 Announce Type: new Abstract: EVICALC, our system for the DocSem shared task, achieved 8.61% joint accuracy on 1,730 tasks in the official final test evaluation. It reads a PDF, selects a passage, asks a language model to write an arithmetic expression, and eval…

  2. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Sachin Gupta ·

    Evidence First, Arithmetic Second: A System Report and Failure Analysis for DocSem

    EVICALC, our system for the DocSem shared task, achieved 8.61% joint accuracy on 1,730 tasks in the official final test evaluation. It reads a PDF, selects a passage, asks a language model to write an arithmetic expression, and evaluates that expression in local code. Saved inter…