A new benchmark called KhatianDoc has been developed to assess the capabilities of multimodal large language models (LLMs) in understanding Bengali legal land records. The benchmark, built from 107 real land records in Bangladesh, includes tasks such as symbol recognition, base-16 to decimal conversion, structured field extraction, and question answering. Evaluations of six LLMs revealed significant failures, with models unable to correctly answer nearly 40% of questions and performing worse than a baseline on arithmetic tasks involving land ownership fractions. AI
IMPACT Highlights critical gaps in multimodal LLM understanding of specialized, non-Latin script data, indicating a need for domain-specific training and evaluation.
RANK_REASON Academic paper introducing a new benchmark for evaluating LLM capabilities on a specific domain. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →