English(EN)Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools
新研究解决多模态推理效率和准确性问题 · 2篇论文
作者PulseAugur 编辑部·[5 个来源]·
两篇新研究论文提出了提高多模态推理模型效率和准确性的方法。第一篇AdaViG引入了一种自适应视觉门控技术,当内部信号表明视觉步骤生成无益时,动态中止该生成,从而提高准确性并降低计算开销。第二篇Beyond the Eye (BEE)专注于自调节隐式视觉工具,将工具调用行为纳入训练目标,以平衡内部知识和外部工具,从而减少延迟和冗余计算。
AI
arXiv:2607.13125v1 Announce Type: cross Abstract: We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image ge…
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editin…
arXiv cs.CV
TIER_1English(EN)·Wenxi Gao, Guanxi Lu, Didi Zhu, Hao Mark Chen, Quan Deng, Zhican Wang, Jiankang Deng, Hongxiang Fan·
arXiv:2607.10004v1 Announce Type: new Abstract: Unified multimodal models (UMMs) with interleaved reasoning, which generate both textual and visual steps as part of intermediate reasoning traces, have demonstrated great potential for visual mathematical reasoning tasks. However, …
arXiv:2607.11106v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have made remarkable progress on fine-grained perception tasks under the "Thinking with Images" (TwI) paradigm by iteratively performing various visual tool operations. However, this p…
Recent multimodal large language models (MLLMs) have made remarkable progress on fine-grained perception tasks under the "Thinking with Images" (TwI) paradigm by iteratively performing various visual tool operations. However, this paradigm relies heavily on frequent external tool…