English(EN)CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering
新基准评估AI代理使用科学软件和执行软件工程任务的能力
作者PulseAugur 编辑部·[5 个来源]·
研究人员为计算机使用代理(CUAs)引入了新的基准和评估框架,这些代理通过图形用户界面进行交互以完成任务。OSWorld-Science 专注于科学软件,包含跨越不同科学领域的 146 个任务,用于测试视觉语言模型(VLMs)。CUA-SWE 针对软件工程任务,要求代理将代码修改与视觉界面交互和验证相结合。此外,OSWorld-Pro 提供了一种基于过程的评估方法,包含超过 2800 个子目标和人工标注,以分析代理的失败模式并提高效率,结果表明即使是 Claude Opus 5 等先进模型在处理这些复杂任务时也面临困难。
AI
arXiv cs.AI
TIER_1English(EN)·Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li, Baiding Chen, Zijun Dou, Qingcheng Zeng, Qi Kang, Oliver Sun, Eric Wang, Bo Zhou, Haixin Wang, Yufan Du, Shi Bo, Ruihan Lin, Mengqi Yuan, Dunjie Lu, Steven Dillmann, Yiming Shi, Tina Su, Amy Xin, Minghao Liu, Xi …·
arXiv:2609.39903v1 Announce Type: new Abstract: Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and produc…
arXiv cs.AI
TIER_1English(EN)·Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh·
arXiv:2609.40284v1 Announce Type: cross Abstract: Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilit…
Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorl…
Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agen…
Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why agen…