Researchers evaluated three open-source vision-language models (VLMs) – InternVL, Qwen-VL, and SmolVLM – on their ability to assess proxemic risk from egocentric robot images. While fine-tuning and advanced prompting strategies showed only modest improvements across the board, Qwen-VL demonstrated a notable increase in correctly identifying high-danger scenarios when using a specific prompt. The study also found that accurate danger classification did not necessarily correlate with better spatial grounding, suggesting that VLMs may provide useful safety labels without precisely attending to the critical areas within an image. AI
IMPACT This research highlights current limitations in VLMs for fine-grained spatial reasoning and safety applications in robotics.
RANK_REASON The cluster describes a research paper evaluating existing models on a specific task.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →