English(EN)WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents
新基准和框架推动多模态AI代理发展
作者PulseAugur 编辑部·[2 个来源]·
研究人员推出了新的框架和基准来改进多模态搜索代理。WeAgent-Harness和WeAgent-MMSearch旨在使代理能够原生交互并引用从网络检索到的图像,解决了当前仅文本方法的局限性。此外,MM-BrowseComp提供了一个包含400个需要视觉证据提取的问题的综合基准,显示即使是先进的模型在多模态浏览方面也面临挑战,准确率仅为24.25%。
AI
arXiv:2608.28062v2 Announce Type: replace Abstract: Multimodal search agents extend parametric knowledge with newly emerging and long-tail evidence from the open web. Yet many existing agentic search environments often expose retrieved evidence only as text and omit tool-returned…
arXiv:2508.13186v2 Announce Type: replace-cross Abstract: AI agents with advanced reasoning and tool-use capabilities have demonstrated impressive performance in web browsing for deep search. However, existing benchmarks such as BrowseComp primarily focus on textual content, over…