Researchers have introduced OSReward, a new benchmark designed to evaluate the reliability of vision-language models (VLMs) when used as judges for computer-using agents (CUAs). The benchmark includes a diverse set of CUA trajectories with ground-truth verdicts derived from multi-stage human annotation. Findings indicate that even state-of-the-art VLMs exhibit a leniency bias, misclassifying failed CUA runs as successful, and the most reliable models are prohibitively expensive for large-scale use. To address this, the team developed OS-Shepherd, a family of open-source reward models trained on a new corpus of reasoning-annotated judgments, offering a cost-effective and stable alternative. AI
IMPACT Establishes a new standard for evaluating AI agents, potentially improving the reliability and cost-effectiveness of their training and deployment.
RANK_REASON The cluster describes a new academic paper introducing a benchmark and models for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]
- computer-using agents
- OSReward
- OSReward-Hard
- OSReward-Multi
- OS-Shepherd
- OS-Shepherd-100K
- vision-language models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →