A new research paper introduces Percept-V, a dataset designed to test the basic visual perception capabilities of multimodal large language models (MLLMs). Despite the simplicity of the tasks, which require minimal reasoning, current state-of-the-art proprietary and open-source MLLMs demonstrated weak performance compared to humans. The study found that model performance significantly degrades as the number of objects in an image increases, and identified specific perception skills that are particularly challenging for these models. While fine-tuning an open-source MLLM showed performance gains, these improvements did not generalize well to related datasets, indicating limitations in learned representations. AI
IMPACT Highlights current limitations in MLLM visual perception, suggesting a need for improved generalization and robustness in foundational models.
RANK_REASON Research paper introducing a new dataset and evaluation methodology for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Hugging Face
- MLLMs
- Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
- Percept-V
- Samrajnee Ghosh
- TVPS-4
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →