A new survey paper published on arXiv examines the evaluation practices for automated research systems. It reviews benchmarks and methodologies across six key areas: literature synthesis, research ideation, workflow execution, scholarly writing, peer review, and end-to-end research. The paper highlights that output checks, process checks, and human studies offer complementary insights, and emphasizes the importance of evaluator calibration and resource budgets for interpreting performance comparisons. Recommendations are provided for reporting and auditing evaluations in specific research settings. AI
IMPACT Provides a framework for evaluating AI-driven research tools, potentially improving their development and adoption.
RANK_REASON The item is a survey paper on benchmarks and evaluation practices for automated research systems. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →