Researchers have introduced Reconstruction, a novel benchmark designed to evaluate language models' ability to infer research ideas solely from pre-publication bibliographies. This benchmark employs strict anti-leakage protocols, including temporal citation cutoffs and anonymous reference IDs, to ensure the integrity of the evaluation. While frontier models achieved modest match rates between 3-15% across six scientific domains and 643 papers, a multi-agent pipeline combining cross-model review with a tournament structure significantly improved performance, reaching match rates of 23-42%. AI
IMPACT This benchmark could drive development of more sophisticated reasoning and information extraction capabilities in LLMs.
RANK_REASON The cluster describes a new benchmark and research paper published on arXiv, detailing a novel evaluation method for language models. [lever_c_demoted from research: ic=1 ai=1.0]
Read on arXiv cs.MA (Multiagent) →
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- ReConStruction
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →