Researchers have introduced new benchmarks and evaluation frameworks for computer-use agents (CUAs), which interact with graphical user interfaces to complete tasks. OSWorld-Science focuses on scientific software, incorporating 146 tasks across various scientific domains to test visual language models (VLMs). CUA-SWE addresses software engineering tasks, requiring agents to integrate code modification with visual interface interaction and verification. Additionally, OSWorld-Pro offers a process-based evaluation method with over 2800 subgoals and human annotations to analyze agent failure modes and improve efficiency, revealing that even advanced models like Claude Opus 5 struggle with these complex tasks. AI
IMPACT These benchmarks will drive progress in developing more capable and efficient AI agents for complex real-world tasks.
RANK_REASON The cluster introduces new academic benchmarks and evaluation frameworks for AI agents, which falls under research.
Read on Hugging Face Daily Papers →
- alphaXiv
- arXiv
- CatalyzeX
- Claude Opus 5
- computer-use agents
- CUA-SWE
- DagsHub
- Gotit.pub
- Hugging Face
- OSWorld-Pro
- OSWorld-Science
- ScienceCast
- visual language models
AI-generated summary · Google Gemini · from 5 sources. How we write summaries →