A new arXiv paper demonstrates that large language models can match human experts in extracting and critically appraising information from scientific publications on microbial oncogenesis. Researchers benchmarked models including GPT-5, GPT-5 Nano, Gemini 2.5 Pro, and Gemini 2.5 Flash using a dataset of 24 research papers on MMTV-LV and breast cancer. GPT-5 and GPT-5 Nano performed indistinguishably from human experts on structured evaluation tasks, suggesting LLMs could be used for automated systematic evidence synthesis. However, the study also identified persistent vulnerabilities in methodological appraisal and contradiction identification within full texts. AI
IMPACT This research suggests LLMs can significantly accelerate scientific literature review and evidence synthesis, potentially speeding up discovery in fields like cancer research.
RANK_REASON The cluster contains an academic paper detailing a new benchmark and evaluation of AI models on a specific research task. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →