ENTITY
OSReward-Hard
OSReward-Hard
PulseAugur coverage of OSReward-Hard — every cluster mentioning OSReward-Hard across labs, papers, and developer communities, ranked by signal.
Total · 30d
2
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
2 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 2 TOTAL
-
New OSReward benchmark reveals VLM judges are too lenient for AI agents
Researchers have introduced OSReward, a new benchmark designed to evaluate the reliability of vision-language models (VLMs) when used as judges for computer-using agents (CUAs). The benchmark includes a diverse set of C…
-
New OSReward benchmark reveals VLM judges are lenient on AI agents
Researchers have introduced OSReward, a new benchmark designed to evaluate the reliability of vision-language models (VLMs) when used as judges for computer-using agents (CUAs). The benchmark includes realistic CUA traj…