A new study published on arXiv evaluates the performance of eight large language models (LLMs) in screening scientific papers for software engineering systematic reviews. The research found that while newer LLMs showed a marginal improvement in performance compared to older models, the differences were not substantial. LLMs still struggle to replace human reviewers for this task, with agreement between models being high but significant disagreements arising on specific inclusion and exclusion criteria. Refining these criteria offered only modest improvements, suggesting future research should focus on agent-based approaches and prompt engineering. AI
IMPACT LLMs show limited progress in automating systematic review screening, indicating human oversight remains crucial.
RANK_REASON Research paper evaluating LLM performance on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- ScienceCast
- scite Smart Citations
- SESR-Eval
- SESR-Eval-Mini
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →