Researchers have introduced DeepSurvey-Bench, a new benchmark designed to evaluate the academic value of automatically generated scientific surveys. Existing benchmarks often rely on superficial criteria like citation counts, which do not fully capture a survey's true academic merit. DeepSurvey-Bench addresses this by incorporating dimensions such as informational, scholarly communication, and research guidance value. The benchmark utilizes a dataset annotated for academic value and employs a multi-LLM-as-a-judge approach to assess generated surveys, demonstrating a closer alignment with human assessments than previous methods. AI
IMPACT Provides a more robust method for evaluating AI-generated academic content, potentially improving future survey generation models.
RANK_REASON The cluster describes a new academic benchmark and associated paper. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →