Researchers have introduced MetaSyn, a new dataset comprising 442 expert-curated meta-analyses from Nature Portfolio journals, designed to benchmark Large Language Model (LLM) agents in scientific reasoning. The dataset includes PI/ECO criteria, a corpus of 140,000 PubMed articles, and verified studies, aiming to evaluate the full pipeline of literature retrieval, study selection, and statistical aggregation. Benchmarking twelve different LLM configurations revealed a significant bottleneck in the screening process, with current systems failing to reliably identify eligible studies from distractors, achieving a maximum recall of only 52.7% despite high retrieval rates. AI
IMPACT This research highlights current limitations in LLM agents for complex scientific reasoning, particularly in study selection, indicating areas for future development.
RANK_REASON The cluster describes a new academic paper introducing a dataset and benchmark for evaluating LLM agents on a specific scientific task.
Read on arXiv cs.IR (Information Retrieval) →
- Hugging Face
- LLM Agents
- Nature Portfolio
- PubMed
- alphaXiv
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Litmaps
- ScienceCast
- scite Smart Citations
- Anzhe Xie
- arXiv
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →