PulseAugur
EN
LIVE 10:26:09

New benchmarks evaluate AI agents' ability to use scientific software and perform software engineering tasks

Researchers have introduced new benchmarks and evaluation frameworks for computer-use agents (CUAs), which interact with graphical user interfaces to complete tasks. OSWorld-Science focuses on scientific software, incorporating 146 tasks across various scientific domains to test visual language models (VLMs). CUA-SWE addresses software engineering tasks, requiring agents to integrate code modification with visual interface interaction and verification. Additionally, OSWorld-Pro offers a process-based evaluation method with over 2800 subgoals and human annotations to analyze agent failure modes and improve efficiency, revealing that even advanced models like Claude Opus 5 struggle with these complex tasks. AI

IMPACT These benchmarks will drive progress in developing more capable and efficient AI agents for complex real-world tasks.

RANK_REASON The cluster introduces new academic benchmarks and evaluation frameworks for AI agents, which falls under research.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 5 sources. How we write summaries →

New benchmarks evaluate AI agents' ability to use scientific software and perform software engineering tasks

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
The cluster introduces new academic benchmarks and evaluation frameworks for AI agents, which falls under research.
Source corroboration
5 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
product, paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
13 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [5]

  1. arXiv cs.AI TIER_1 English(EN) · Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li, Baiding Chen, Zijun Dou, Qingcheng Zeng, Qi Kang, Oliver Sun, Eric Wang, Bo Zhou, Haixin Wang, Yufan Du, Shi Bo, Ruihan Lin, Mengqi Yuan, Dunjie Lu, Steven Dillmann, Yiming Shi, Tina Su, Amy Xin, Minghao Liu, Xi … ·

    OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

    arXiv:2609.39903v1 Announce Type: new Abstract: Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and produc…

  2. arXiv cs.AI TIER_1 English(EN) · Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh ·

    cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

    arXiv:2609.40284v1 Announce Type: cross Abstract: Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilit…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

    Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorl…

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering

    Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agen…

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    OSWorld-Pro: Process-based Evaluation for Computer Use Agents

    Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why agen…