PulseAugur
实时 16:35:08
English(EN) DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection

新AI研究聚焦多模态推理、效率和机器人感知

arXiv上发布的几篇研究论文提出了改进AI模型多模态推理的新方法。VISE(Visual Invariance Self-Evolution)通过强制执行空间和语义不变性来解决视觉欠条件问题,在图像字幕和VQA任务上取得了显著的提升。Visual-OPSD专注于高效推理,通过将使用特权视觉思维的教师模型的知识蒸馏到一个纯文本学生模型中,实现了显著的加速。另一种方法Ask, Solve, Generate,使用自我一致性奖励在没有外部监督的情况下自主改进视觉理解和图像生成。Position Rebinding Cache Reuse (PRCR)解决了视觉缓存中过时的位置绑定问题,实现了无重放的视觉重访并减少了计算量。最后,OctoSense提出了一个使用多种传感器进行多模态机器人感知的自监督学习框架,在各种任务上的表现优于仅图像模型。 AI

影响 这些论文引入了改进多模态推理、效率和机器人感知的新技术,有可能提升AI系统在复杂任务中的能力。

排序理由 arXiv上发表的多篇研究论文详细介绍了多模态AI的新方法。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 14 个来源。 我们如何撰写摘要 →

新AI研究聚焦多模态推理、效率和机器人感知

报道来源 [14]

  1. arXiv cs.AI TIER_1 English(EN) · Yujin Tang, Chenming Shang, Ruize Xu, Nikhil Singh ·

    DMV-Bench:通过附带线索注入诊断长时域多模态代理的视觉记忆

    arXiv:2606.27499v1 Announce Type: cross Abstract: Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when an agent genuinely needs to remember what it saw rather than what it could write …

  2. arXiv cs.CL TIER_1 English(EN) · Nikhil Singh ·

    DMV-Bench:通过附带线索注入诊断长视域多模态智能体的视觉记忆

    Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when an agent genuinely needs to remember what it saw rather than what it could write down. We introduce DMV-Bench (Code: https://github…

  3. arXiv cs.CV TIER_1 English(EN) · Mengzhao Wang, Yanli Ji, Wangmeng Zuo, Peng Ye, Chongjun Tu ·

    Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning

    arXiv:2606.26631v1 Announce Type: new Abstract: Interleaved multimodal reasoning improves visual grounding by revisiting visual evidence during multi-step generation, yet existing methods typically rely on token replay, repeatedly forwarding selected visual tokens. A natural shor…

  4. arXiv cs.CV TIER_1 English(EN) · Anthony Bisulco, Jeremy Wang, Kostas Daniilidis, Randall Balestriero, Pratik Chaudhari ·

    OctoSense:用于多模态机器人感知的自监督学习

    arXiv:2606.27317v1 Announce Type: new Abstract: We present OctoSense, an open-source sensor platform with stereo RGB and event cameras, LiDAR, a thermal camera, an inertial measurement unit, RTK-corrected global positioning system, and proprioception (CAN bus data from a car, and…

  5. arXiv cs.CV TIER_1 English(EN) · Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Salman Khan, Fahad Khan ·

    在自进化大型多模态模型中更加关注视觉Token

    arXiv:2606.27373v1 Announce Type: new Abstract: Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and self-consistency reward schemes in existing self-evolv…

  6. arXiv cs.CV TIER_1 English(EN) · Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar, Abdelrahman Shaker, Fahad Khan, Hisham Cholakkal, Salman Khan, Rao Muhammad Anwer ·

    提问、解决、生成:通过自洽性奖励实现自进化的统一多模态理解与生成

    arXiv:2606.27376v1 Announce Type: new Abstract: Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward models. We ask…

  7. arXiv cs.CV TIER_1 English(EN) · Pengyu Li, Zhitao Gao, Lingling Zhang, Muye Huang, Yuanming Li, Fangzhi Xu, Jun Liu ·

    Visual-OPSD:面向高效统一多模态推理的跨模态 on-policy 自蒸馏

    arXiv:2606.18974v2 Announce Type: replace Abstract: Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly an order-of-magnitude inference cost from multi-step diffusion. We find this cost …

  8. arXiv cs.CV TIER_1 English(EN) · Rao Muhammad Anwer ·

    提问、解决、生成:通过自洽性奖励实现自演化统一多模态理解与生成

    Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward models. We ask whether a unified LMM can improve both abilitie…

  9. arXiv cs.CV TIER_1 English(EN) · Fahad Khan ·

    在自进化大型多模态模型中更加关注视觉Token

    Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and self-consistency reward schemes in existing self-evolving LMMs optimize answer agreement without ensur…

  10. arXiv cs.CV TIER_1 English(EN) · Pratik Chaudhari ·

    OctoSense:用于多模态机器人感知的自监督学习

    We present OctoSense, an open-source sensor platform with stereo RGB and event cameras, LiDAR, a thermal camera, an inertial measurement unit, RTK-corrected global positioning system, and proprioception (CAN bus data from a car, and joint angles for a quadruped robot). The eponym…

  11. arXiv cs.CV TIER_1 English(EN) · Chongjun Tu ·

    位置重绑定缓存复用:无重放交错多模态推理视觉重访

    Interleaved multimodal reasoning improves visual grounding by revisiting visual evidence during multi-step generation, yet existing methods typically rely on token replay, repeatedly forwarding selected visual tokens. A natural shortcut is to reuse the historical visual key-value…

  12. arXiv cs.CV TIER_1 English(EN) · Xiuwei Chen, Wentao Hu, Yongxin Wang, Zisheng Chen, Likui Zhang, Kun Xiang, Jianhua Han, Hui-Ling Zhen, Jingyuan Zou, Hang Xu, Xiaodan Liang ·

    用于高效多模态推理的潜在视觉状态

    arXiv:2606.24233v1 Announce Type: new Abstract: The integration of visual evidence has significantly enhanced the capabilities of large multimodal models. However, this integration predominantly relies on generating discrete outputs (etc., code or box coordinates) to invoke exter…

  13. arXiv cs.CV TIER_1 English(EN) · Xiaodan Liang ·

    用于高效多模态推理的潜在视觉状态

    The integration of visual evidence has significantly enhanced the capabilities of large multimodal models. However, this integration predominantly relies on generating discrete outputs (etc., code or box coordinates) to invoke external tools, a process that introduces rigid depen…

  14. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    多模态AI:一个能看、能读、能听的模型

    <p>The models you use now don't just read — they see, hear, and read images too. GPT-4o, Gemini, Claude with vision: all multimodal. The trick that makes it work is the same embeddings idea, stretched across senses. Here's how, visualized.</p> <p>👁️‍🗨️ <strong>Watch modalities me…