A new framework called DocAnnot has been developed to accelerate the creation of datasets for Key Information Extraction (KIE). DocAnnot utilizes a Large Vision Language Model (LVLM) for extracting label values, combined with OCR for text and bounding box detection. A novel Spatially Informed Contextual Matching (SICM) algorithm enhances label-value association by considering spatial relationships and proximity alongside textual data. While models trained solely on DocAnnot's auto-generated data achieve respectable performance, human-annotated data remains superior, though DocAnnot significantly reduces manual effort and costs. AI
IMPACT Streamlines KIE dataset creation, potentially reducing costs and accelerating model development for document analysis tasks.
RANK_REASON The cluster contains a research paper detailing a new framework and algorithm for dataset creation. [lever_c_demoted from research: ic=1 ai=1.0]
- CORD
- DocAnnot
- generative artificial intelligence
- Harikrishnan P M Dr
- LayoutLMv3
- optical character recognition
- Spatially Informed Contextual Matching
- SROIE
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →