New research indicates that vision-language models (VLMs) struggle with affordance prediction, primarily due to difficulties in correctly identifying object parts rather than a lack of action knowledge. Studies using benchmarks like GroundBench show that while naming the target part significantly improves a model's ability to predict the correct action, some models may rely on textual shortcuts rather than genuine visual grounding. This suggests that improving part identification is key to enhancing VLM performance in robotic manipulation tasks. AI
IMPACT Highlights a key bottleneck in VLM affordance prediction, suggesting improvements in part grounding could significantly enhance robotic manipulation capabilities.
RANK_REASON The cluster contains two academic papers detailing new benchmarks and findings related to vision-language models and their performance on affordance prediction tasks.
- arXiv
- button
- Camera
- GPT-4o mini
- GPT-5
- GroundBench
- Hugging Face
- OpenAI
- Push
- robot
- vision-language model
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →