PulseAugur
实时 02:17:41
English(EN) PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking

新框架解耦 AI 感知与推理,增强视觉理解能力 · 追踪 6 个来源

研究人员推出了新颖的框架,以增强视觉语言模型细粒度的视觉推理能力。Rule-VLN 通过引入大规模城市基准和语义导航纠正模块 (SNRM) 来灌输安全意识,解决了具身 AI 代理优先考虑物理导航而非语义规则的挑战。此外,Perceive-to-Reason (P2R) 和 PixelEyes 提出了将感知与推理解耦的方法,提高了在高分辨率图像上的性能,并减少了多轮视觉推理任务中的冗余轨迹。 AI

影响 感知与推理解耦方面的这些进步可能导致在现实世界应用中,尤其是在导航和复杂视觉任务中,AI 代理更加鲁棒和合规。

排序理由 该集群包含多篇研究论文,介绍了 AI 视觉推理的新基准、模块和框架。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →

新框架解耦 AI 感知与推理,增强视觉理解能力 · 追踪 6 个来源

报道来源 [6]

  1. arXiv cs.AI TIER_1 English(EN) · Jiawen Wen, Penglei Sun, Wenjie Zhang, Suixuan Qiu, Weisheng Xu, Xiaofei Yang, Xiaowen Chu ·

    Rule-VLN:通过语义推理和几何校正实现感知与合规的桥梁

    arXiv:2604.16993v2 Announce Type: replace Abstract: As embodied AI transitions to real-world deployment, the success of the Vision-and-Language Navigation (VLN) task tends to evolve from mere reachability to social compliance. However, current agents suffer from a "goal-driven tr…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    Perceive-to-Reason:解耦感知与推理,实现细粒度视觉推理

    A unified framework named Perceive-to-Reason (P2R) is introduced that separates visual perception from reasoning in vision-language models through a two-stage process, improving fine-grained visual reasoning performance on high-resolution images.

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    PixelEyes:解耦感知与推理,实现精准视觉证据搜寻

    Multi-turn visual reasoning agents suffer from entangled reasoning and perception that cause redundant trajectories; PixelEyes addresses this by decoupling these processes through mask-guided search and semantic-region breadth-first search, demonstrated on a new benchmark with ex…

  4. arXiv cs.CV TIER_1 English(EN) · Dengxian Gong, Yuanzheng Wu, Haobo Yuan, Zhengdong Hu, Tao Zhang, Yikang Zhou, Shihao Chen, Quanzhu Niu, Kai Wang, Jason Li, Haochen Wang, Lu Qi, Shunping Ji, Ming-Hsuan Yang ·

    PixelEyes:解耦感知与推理,实现精准视觉证据搜寻

    arXiv:2607.00115v1 Announce Type: new Abstract: This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leading to long, redundant trajectories. We attribute this failure to the entanglement of reasoning and perception withi…

  5. arXiv cs.CV TIER_1 English(EN) · Hongxing Li, Xiufeng Huang, Dingming Li, Wenjing Jiang, Zixuan Wang, Haolei Xu, Hanrong Zhang, Haiwen Hong, Longtao Huang, Hui Xue, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen ·

    感知到推理:解耦感知与推理以实现细粒度视觉推理

    arXiv:2607.01191v1 Announce Type: new Abstract: Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visual cues are buried in high-resolution images. Existing approaches rely on repeated cropping or test-time visual sea…

  6. arXiv cs.CV TIER_1 English(EN) · Yongliang Shen ·

    Perceive-to-Reason:解耦感知与推理,实现细粒度视觉推理

    Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visual cues are buried in high-resolution images. Existing approaches rely on repeated cropping or test-time visual search to introduce local evidence, but they typica…