Researchers have introduced Khondo, a new benchmark designed to evaluate multimodal large language models (MLLMs) on the task of splitting document packets into their constituent parts. This benchmark is unique as it focuses on Bangla government forms, is bilingual (Bangla-English), and operates directly on page images rather than OCR text. Initial evaluations show that MLLMs can effectively cluster pages belonging to the same document but struggle significantly with reconstructing the original page order, especially for shuffled Bangla documents. The study highlights that explicit instructions are crucial for ordering, and English packets are handled better than Bangla, indicating that page arrangement is a primary challenge, with language being a secondary but consistent factor. AI
IMPACT Establishes a new benchmark for low-resource document understanding, pushing MLLMs to improve page ordering capabilities.
RANK_REASON The item describes a new academic benchmark for evaluating multimodal large language models on a specific document understanding task. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →