Researchers have introduced SciLitBench, a new benchmark designed to evaluate the capabilities of large language models (LLMs) in performing systematic literature reviews. The benchmark covers multiple stages, including title and abstract screening, full-text screening, and data extraction, utilizing a dataset of over 42,000 records. Experiments with 22 open-weight LLMs revealed that while explicit criteria improve screening performance, data extraction accuracy varies significantly, with models struggling to accurately extract detailed information like computational approaches or limitations. SciLitBench highlights a current limitation in LLM performance for comprehensive evidence synthesis, differentiating between high-recall screening and detailed data extraction. AI
IMPACT Identifies practical limitations of current LLMs in complex evidence synthesis tasks, guiding future research and development.
RANK_REASON The cluster describes a new benchmark and research paper evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Litmaps
- LLMs
- ScienceCast
- SciLitBench
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →