A new study explores the use of Large Language Models (LLMs) for conducting systematic literature reviews (SLRs) in the field of disease spread modeling. Researchers developed an LLM pipeline to extract information from 536 agent-based modeling papers, comparing its performance against a human-conducted SLR. The study found that GPT-4.1 achieved approximately 77.95% paper-level accuracy, while GPT-5.0 reached 81.67%, with field-level accuracies varying significantly. Notably, the agreement between LLMs was identified as a potential indicator of output quality, with low agreement suggesting hallucinations and high agreement with low accuracy pointing to noise in the human dataset. AI
IMPACT LLMs can potentially automate and improve the efficiency of research processes like systematic literature reviews in specialized scientific domains.
RANK_REASON The cluster describes an academic paper detailing a new methodology using LLMs for systematic literature reviews. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →