PulseAugur
中
实时 11:37:31
English(EN) CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering

新基准评估AI代理使用科学软件和执行软件工程任务的能力

研究人员为计算机使用代理(CUAs)引入了新的基准和评估框架,这些代理通过图形用户界面进行交互以完成任务。OSWorld-Science 专注于科学软件,包含跨越不同科学领域的 146 个任务,用于测试视觉语言模型(VLMs)。CUA-SWE 针对软件工程任务,要求代理将代码修改与视觉界面交互和验证相结合。此外,OSWorld-Pro 提供了一种基于过程的评估方法,包含超过 2800 个子目标和人工标注,以分析代理的失败模式并提高效率,结果表明即使是 Claude Opus 5 等先进模型在处理这些复杂任务时也面临困难。 AI

影响 这些基准将推动开发更强大、更高效的 AI 代理以应对复杂的现实世界任务。

排序理由 该集群引入了新的学术基准和 AI 代理评估框架,属于研究范畴。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

新基准评估AI代理使用科学软件和执行软件工程任务的能力

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群引入了新的学术基准和 AI 代理评估框架,属于研究范畴。
Source corroboration
5 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
product, paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
13 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [5]

  1. arXiv cs.AI TIER_1 English(EN) · Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li, Baiding Chen, Zijun Dou, Qingcheng Zeng, Qi Kang, Oliver Sun, Eric Wang, Bo Zhou, Haixin Wang, Yufan Du, Shi Bo, Ruihan Lin, Mengqi Yuan, Dunjie Lu, Steven Dillmann, Yiming Shi, Tina Su, Amy Xin, Minghao Liu, Xi … ·

    OSWorld-科学:用于学习和使用科学软件的计算机使用代理基准测试

    arXiv:2609.39903v1 Announce Type: new Abstract: Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and produc…

  2. arXiv cs.AI TIER_1 English(EN) · Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh ·

    cua-speedrun: 计算机使用代理速度标准化基准测试

    arXiv:2609.40284v1 Announce Type: cross Abstract: Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilit…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    OSWorld-科学:用于学习和使用科学软件的计算机使用代理基准测试

    Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorl…

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    CUA-SWE:当计算机使用代理遇上视觉软件工程

    Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agen…

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    OSWorld-Pro:面向计算机使用代理的基于过程的评估

    Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why agen…