OSReward
PulseAugur coverage of OSReward — every cluster mentioning OSReward across labs, papers, and developer communities, ranked by signal.
-
AI agent grading flawed: VLM judges prioritize narration over action
A new analysis by OSReward reveals a flaw in how computer-use agents are evaluated, specifically concerning the use of VLM judges. These judges appear to prioritize an agent's self-reported completion messages over the …
-
AI agent evaluation models show bias, inflating success rates
Recent discussions about AI hype cycles are being challenged by a closer examination of evaluation methods. The OSReward project highlights that reward models used to judge AI agents are not only noisy but also biased, …
-
New OSReward benchmark reveals VLM judges are too lenient for AI agents
Researchers have introduced OSReward, a new benchmark designed to evaluate the reliability of vision-language models (VLMs) when used as judges for computer-using agents (CUAs). The benchmark includes a diverse set of C…
-
New OSReward benchmark reveals VLM judges are lenient on AI agents
Researchers have introduced OSReward, a new benchmark designed to evaluate the reliability of vision-language models (VLMs) when used as judges for computer-using agents (CUAs). The benchmark includes realistic CUA traj…