A new research paper titled "Composition, Not Conversation: VLMs Lose the Scene, Not the Thread" explores the limitations of Vision-Language Models (VLMs) when presented with fragmented visual information. The study introduces Layered-VQA, a dataset of 93 scenes and 300 questions, where images are decomposed into RGBA layers. Experiments with eleven open-weight VLMs and two proprietary models revealed consistent failures in composition, grounding, and evidence utilization. The findings indicate that while fragmenting questions has a minor impact, decomposing scenes significantly reduces VLM accuracy, suggesting that the way visual evidence is presented is crucial for model performance. AI
IMPACT Highlights critical limitations in VLM scene composition and evidence grounding, suggesting a need for new evaluation methods.
RANK_REASON Research paper published on arXiv detailing limitations of Vision-Language Models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Layered-VQA
- Lekkala Sai Teja
- ScienceCast
- Vision--Language Models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →