A system named EVICALC, designed for the DocSem shared task, achieved 8.61% joint accuracy on 1,730 tasks. The system processes PDFs by selecting passages, using a language model to generate arithmetic expressions, and then evaluating these expressions locally. A separate public-validation run yielded higher scores of 92.17% answer accuracy and 1.00 evidence F1, though these metrics are not directly comparable due to differing configurations. Analysis revealed that issues such as optical character recognition errors merging text blocks can lead to the system answering from unrelated content. AI
IMPACT Highlights challenges in integrating OCR and LLMs for complex document understanding tasks.
RANK_REASON The cluster contains a research paper detailing a system's performance and failure analysis on a specific task.
Read on arXiv cs.IR (Information Retrieval) →
- alphaXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- DocSem
- EVICALC
- Gotit.pub
- Hugging Face
- Litmaps
- optical character recognition
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →