PulseAugur
实时 10:11:37
English(EN) SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance

SpatialAfford框架增强了紧凑型VLM的精确可供性定位能力

研究人员开发了SpatialAfford,一个新颖的两阶段框架,旨在提高紧凑型视觉语言模型(VLM)的可供性定位能力。该框架首先使用空间注意力对齐(SAA)将模型的注意力引导到正确的可供性区域,然后使用空间感知GRPO进行坐标预测的精炼。通过在模型生成坐标之前明确引导模型的视觉焦点,SpatialAfford增强了空间推理能力,适用于抓取或按压特定物体部件等任务。在ShareRobot-Bench、ReasonAff和PartAfford等基准测试上的实验表明,SpatialAfford持续提升性能,其中一个较小的4B模型超越了更大的7B+基线模型。 AI

影响 通过增强紧凑型VLM识别特定交互区域的能力,提高了它们在具身AI任务中的精确度。

排序理由 研究论文,详细介绍了一个用于改进视觉语言模型的新框架。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

SpatialAfford框架增强了紧凑型VLM的精确可供性定位能力

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Yufei Zhang, Chenlu Zhan, Donghui Sun, Xiaoxin Chen, Hongwei Wang ·

    SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance

    arXiv:2608.00502v1 Announce Type: new Abstract: Affordance grounding aims to localize the functional region for interaction, such as the handle to grasp or the button to press, rather than the whole object. This makes it more challenging than generic visual grounding because the …