PulseAugur
实时 08:11:35
English(EN) ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision

AI模型通过文本脚手架学习视觉工具使用,而非仅像素·跟踪3篇论文

三篇新研究论文探讨了视觉工具在多模态AI模型中的使用效果。ToolVision和OpenVisTool提出了新的方法,通过关注工具生成信息的因果效用和必要性来训练模型使用视觉工具,而不是仅仅模仿工具调用模式。第三篇论文TextCall提出一个假设,即工具调用的文本脚手架,包括其意图和参数,比返回的像素本身对视觉推理更为关键。这些方法旨在改进AI模型学习和利用视觉证据以获得更好推理能力的方式。 AI

影响 这些论文表明,关注工具调用的文本脚手架,而不仅仅是返回的像素,可以带来更有效和高效的AI模型视觉推理能力。

排序理由 三篇在arXiv上发表的学术论文,详细介绍了AI模型视觉工具使用的新方法和假设。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

AI模型通过文本脚手架学习视觉工具使用,而非仅像素·跟踪3篇论文

报道来源 [3]

  1. arXiv cs.AI TIER_1 English(EN) · Delin Mao, Chenghao Sun, Jingwei Song, Chishui Chen, Linfeng Zhang ·

    ToolVision:学习何时以及如何使用视觉工具进行能力对齐的监督

    arXiv:2608.08907v1 Announce Type: cross Abstract: Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevailing SFT-then-RL recipe creates a different supervision misalignment at each stage. SFT is ex…

  2. arXiv cs.CL TIER_1 English(EN) · Changhao Xiang, Shilin Zhang, Zheng Ma, Kanzhi Cheng, Ruize Ma, Yi Feng, Jianbing Zhang, Zhi Wang, Zhen Wu, Xinyu Dai, Lewei Lu ·

    OpenVisTool:合成教学性视觉工具使用轨迹的开放式配方

    arXiv:2608.08557v1 Announce Type: new Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for …

  3. arXiv cs.CV TIER_1 English(EN) · Jiahao Shao, Yuanbo Yang, Yiyi Liao, Yujun Shen, Ceyuan Yang, Yinghao Xu ·

    用工具思考,而非像素:工具调用作为视觉推理的文本脚手架

    arXiv:2608.09682v1 Announce Type: new Abstract: Tool-augmented vision-language models increasingly "think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind tests, gain decompositions, and attention analyses has…