Researchers have introduced TomaMMU, a large-scale dataset for understanding tomato leaf diseases, and TomaBench, a benchmark designed to evaluate Vision-Language Models (VLMs) on this task. The dataset includes over 28,000 images and more than 200,000 annotated question-answer pairs, structured across seven agricultural tasks. Initial evaluations revealed that current state-of-the-art VLMs struggle with fine-grained recognition and factually grounded reasoning in this domain, though simple fine-tuning on TomaMMU significantly improved performance. AI
IMPACT This benchmark could drive improvements in VLM capabilities for specialized agricultural diagnostics.
RANK_REASON The cluster describes a new academic paper introducing a dataset and benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- computer science
- Computer vision and pattern recognition
- Hugging Face
- Khang Nguyen Quoc
- TomaBench
- TomaMMU
- Vision--Language Models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →