PulseAugur
实时 05:45:23
English(EN) Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?

视觉语言模型用于机器人安全风险评估

研究人员评估了三种开源视觉语言模型(VLM)——InternVL、Qwen-VL 和 SmolVLM——在评估机器人第一人称图像近邻风险方面的能力。尽管微调和高级提示策略在整体上仅显示出适度改进,但 Qwen-VL 在使用特定提示时,在正确识别高危险场景方面表现出显著的提高。研究还发现,准确的危险分类并不一定与更好的空间定位相关,这表明 VLM 可以在不精确关注图像关键区域的情况下提供有用的安全标签。 AI

影响 这项研究突显了 VLM 在机器人精细空间推理和安全应用方面的现有局限性。

排序理由 该集群描述了一篇评估现有模型在特定任务上表现的研究论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

视觉语言模型用于机器人安全风险评估

报道来源 [2]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    视觉语言模型能否评估以自我为中心的机器人图像的邻近风险?

    Assessing proxemic danger from a robot's egocentric perspective is critical for safe embodied navigation in human environments and requires both visual and contextual reasoning. We evaluate three opensource vision-language models (VLMs) (\textit{InternVL}, \textit{Qwen-VL}, and \…

  2. arXiv cs.CV TIER_1 English(EN) · Vladyslava Rudas, Dmytro Kuzmenko ·

    视觉-语言模型能否评估以自我为中心的机器人图像中的空间距离风险?

    arXiv:2608.12515v1 Announce Type: new Abstract: Assessing proxemic danger from a robot's egocentric perspective is critical for safe embodied navigation in human environments and requires both visual and contextual reasoning. We evaluate three opensource vision-language models (V…