Researchers have developed a new benchmark called Research Attention Prediction (RAP) to evaluate how well large language models can track shifts in research attention within the AI/ML field. The benchmark, covering 278 AI/ML fields and 1,390 episodes, involves LLM agents searching a restricted arXiv corpus to predict paper shares over the next six months. Results indicate that while search generally helps, models often underperform a simple exponential moving average baseline, highlighting bottlenecks in evidence acquisition and future-specific updating. Fine-tuning on realized outcomes showed improvements for specific models like Qwen3-4B. AI
IMPACT This benchmark could lead to more sophisticated AI research agents capable of understanding and predicting scientific trends.
RANK_REASON The cluster contains a research paper detailing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →