Researchers evaluated three open-source vision-language models (VLMs) – InternVL, Qwen-VL, and SmolVLM – on their ability to assess proxemic risk from egocentric robot images. While all models performed near a baseline without fine-tuning, Qwen-VL demonstrated superior recall for high-danger cases when combined with an advanced prompting strategy and QLoRA fine-tuning. The study also found that accurate danger classification did not correlate with better spatial grounding, suggesting current VLMs have limitations in fine-grained proxemic reasoning and spatial awareness. AI
IMPACT This research highlights limitations in current VLMs for real-world robot navigation safety, suggesting areas for future development in spatial reasoning and grounding.
RANK_REASON The cluster contains an academic paper detailing research on vision-language models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →