A new pilot study has evaluated the capabilities of multimodal large language models (MLLMs) in understanding low-resource Khmer documents. Researchers found that while current MLLMs can process visually clear English and structured numeric content, reliable native Khmer document understanding remains a significant challenge. The study constructed an evaluation subset from the KH-FUNSD collection, testing Qwen-VL models and finding that external OCR tools like Tesseract and PaddleOCR yielded better results than direct prompting for Khmer-script answers. AI
IMPACT Current MLLMs show limitations in processing non-Latin scripts and mixed-language documents, indicating a need for further development in low-resource language understanding.
RANK_REASON Academic paper presenting a pilot study and evaluation of MLLMs on a specific low-resource language document set. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →