Researchers have introduced FlyAOC, a new benchmark designed to evaluate AI agents on the complex task of curating scientific knowledge bases. This benchmark simulates the end-to-end workflow of expert curators, requiring agents to search through scientific literature, reconcile evidence, and produce structured annotations. FlyAOC uses a corpus of 16,898 papers and 7,397 expert-curated annotations across 100 genes from FlyBase, the Drosophila knowledge base, to assess agents' ability to recover standardized function terms, expression patterns, and historical synonyms. Initial evaluations using various agent harnesses revealed that system-level failure modes, not apparent in model-only evaluations, are detectable with FlyAOC. AI
IMPACT Provides a new evaluation framework for AI agents in scientific knowledge curation, potentially accelerating discovery.
RANK_REASON The item is a research paper introducing a new benchmark for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Drosophila
- FlyAOC
- FlyBase
- Gotit.pub
- Hugging Face
- ScienceCast
- Xingjian Zhang
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →