PulseAugur
EN
LIVE 13:36:31

New Percept-V dataset reveals MLLMs struggle with basic visual perception

A new research paper introduces Percept-V, a dataset designed to test the basic visual perception capabilities of multimodal large language models (MLLMs). Despite the simplicity of the tasks, which require minimal reasoning, current state-of-the-art proprietary and open-source MLLMs demonstrated weak performance compared to humans. The study found that model performance significantly degrades as the number of objects in an image increases, and identified specific perception skills that are particularly challenging for these models. While fine-tuning an open-source MLLM showed performance gains, these improvements did not generalize well to related datasets, indicating limitations in learned representations. AI

IMPACT Highlights current limitations in MLLM visual perception, suggesting a need for improved generalization and robustness in foundational models.

RANK_REASON Research paper introducing a new dataset and evaluation methodology for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New Percept-V dataset reveals MLLMs struggle with basic visual perception

How we ranked this

Signal score
7 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Research paper introducing a new dataset and evaluation methodology for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

Full methodology in our editorial standards.

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Samrajnee Ghosh, Ashish Goswami, Naman Agarwal, Hemanshu Garg, Chinmay Mittal, Mausam, Parag Singla ·

    The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?

    arXiv:2508.21143v4 Announce Type: replace Abstract: Cognitive science research treats visual perception, the ability to understand and make sense of a visual input, as one of the early developmental signs of intelligence. Its TVPS-4 framework categorizes and tests human perceptio…