Researchers have introduced ChitraMiti-12.8k, a new benchmark designed to evaluate visual grounding and modality reliance in Bengali geometric reasoning for vision-language models (VLMs). The benchmark includes synthetic geometry problems and complementary diagrams from school textbooks. Evaluations across several VLMs revealed that while structured descriptions can serve as a proxy for visual input, models struggle with cross-modal verification and are easily misled by textual inaccuracies, even when answering correctly. Fine-tuning on ChitraMiti-12.8k shows improvement, but a significant performance gap persists compared to the strongest zero-shot models. AI
IMPACT This benchmark could advance the evaluation of multimodal reasoning in low-resource languages, pushing VLM development towards more robust cross-modal verification.
RANK_REASON The cluster contains an academic paper introducing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →