PulseAugur
EN
LIVE 22:32:23

Vision-language models lack agency and knowledge retention, new papers reveal

Two new research papers highlight limitations in current vision-language models (VLMs), particularly concerning their ability to retain knowledge after fine-tuning and their lack of "agency" in visual reasoning. The first paper, "Does VLA Even Know the Basics?", introduces Act2Answer, a protocol to evaluate embodied VLA models by having them select answers through actions, revealing that while they perform well on simple concepts, they struggle with richer semantic categories compared to their source VLMs. The second paper, "Position: The Systemic Lack of Agency in Visual Reasoning", argues that VLMs are constrained by a lack of agency, leading them to act as passive semantic retrievers rather than active explorers of visual information, a gap addressed by their proposed Visual Implicit Reasoning Diagnosing Benchmark (V-IRD). AI

IMPACT Highlights critical gaps in current vision-language models, suggesting a need for new evaluation methods and architectures that foster active reasoning and better knowledge retention.

RANK_REASON Two academic papers published on arXiv discussing limitations of vision-language models.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Vision-language models lack agency and knowledge retention, new papers reveal

COVERAGE [1]

  1. arXiv cs.CV TIER_1 English(EN) · Yizhao Huang, Haoyang Chen, Shiqin Wang, Pohsun Huang, Jiayuan Li, Haoyuan Du, Yandong Shi, Zheng Wang, Zhixiang Wang ·

    Position: The Systemic Lack of Agency in Visual Reasoning

    arXiv:2606.14795v1 Announce Type: new Abstract: This paper argues that a systemic lack of Agency constrains the implicit reasoning capabilities of current Vision-Language Models (VLMs). Implicit reasoning refers to the ability to autonomously discover and utilize hidden visual ev…