A new study published on arXiv evaluates the effectiveness of Large Language Models (LLMs) in screening research papers for evidence synthesis. The research found that while no workflow, human or LLM, could identify all relevant studies, LLM performance varied significantly based on the processing configuration rather than the model itself. The study suggests that LLMs are best suited for validated, human-supervised workflows in high-recall scenarios, rather than autonomous exclusion, due to issues like run-to-run consistency. AI
IMPACT LLM screening performance is highly dependent on workflow configuration, suggesting a need for careful integration into human-supervised research processes.
RANK_REASON The cluster contains an academic paper detailing research findings on LLM performance. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →