PulseAugur
EN
LIVE 00:04:46

New benchmark FlyAOC evaluates AI agents on scientific knowledge curation

Researchers have introduced FlyAOC, a new benchmark designed to evaluate AI agents on the complex task of curating scientific knowledge bases. This benchmark simulates the end-to-end workflow of expert curators, requiring agents to search through scientific literature, reconcile evidence, and produce structured annotations. FlyAOC uses a corpus of 16,898 papers and 7,397 expert-curated annotations across 100 genes from FlyBase, the Drosophila knowledge base, to assess agents' ability to recover standardized function terms, expression patterns, and historical synonyms. Initial evaluations using various agent harnesses revealed that system-level failure modes, not apparent in model-only evaluations, are detectable with FlyAOC. AI

IMPACT Provides a new evaluation framework for AI agents in scientific knowledge curation, potentially accelerating discovery.

RANK_REASON The item is a research paper introducing a new benchmark for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark FlyAOC evaluates AI agents on scientific knowledge curation

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Xingjian Zhang, Sophia Moylan, Ziyang Xiong, Qiaozhu Mei, Yichen Luo, Jiaqi W. Ma ·

    FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases

    arXiv:2602.09163v2 Announce Type: replace Abstract: Scientific knowledge bases accelerate discovery by curating findings from primary literature into structured, queryable formats for both human researchers and emerging AI systems. Maintaining these resources requires expert cura…