The new Team Recruitment (Oracle) benchmark, developed by lforla, evaluates AI agents on their ability to query a résumé oracle and assemble a team that meets specific criteria. Currently, the leaderboard features two free-tier models from opencode-zen: Nemotron 3 Ultra, which achieved a score of 90.87, and HY3, with a score of 83.1. This benchmark differs from static knowledge tests by simulating a dynamic workflow that requires agents to plan queries, interpret structured data, and make decisions under constraints, aiming to assess performance based on outcome quality. AI
IMPACT Establishes a new evaluation method for AI agents, focusing on dynamic workflows and outcome quality rather than static knowledge recall.
RANK_REASON New benchmark and leaderboard release for AI agent evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →