Researchers have introduced TomaMMU, a large-scale dataset for multimodal understanding of tomato leaf diseases, and TomaBench, a benchmark designed to evaluate Vision-Language Models (VLMs) on these tasks. The dataset includes over 28,000 images and 213,000 annotated question-answer pairs, organized into seven agricultural tasks across three levels of complexity, from basic perception to expert diagnosis. Evaluations of 14 state-of-the-art VLMs revealed significant gaps in fine-grained recognition and reasoning, though simple fine-tuning on TomaMMU substantially improved performance. AI
IMPACT This benchmark could drive improvements in specialized VLM capabilities for agricultural diagnostics and fine-grained visual understanding.
RANK_REASON The cluster describes a new academic paper introducing a dataset and benchmark for evaluating multimodal AI models. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
- Hugging Face
- Munich Center for Quantum Science and Technology
- TomaBench
- TomaMMU
- Vision--Language Models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →