Researchers have developed NovGauge, a new benchmark designed to fine-tune and diagnose the capabilities of large language models (LLMs) in assessing the novelty of research papers. The benchmark, comprising 619 paper pairs and 50 multi-paper sets, evaluates novelty across three dimensions: task, problem, and method. Initial evaluations of 18 LLMs revealed significant issues with hallucination rates and evidence faithfulness, with even the top-performing model, GPT-5.5, struggling to maintain accuracy after verification. AI
IMPACT Highlights the limitations of current LLMs in scientific literature analysis, indicating a need for further development in reasoning and evidence grounding.
RANK_REASON The item describes a new benchmark for evaluating LLMs, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- GPT-5.5
- Hugging Face
- International Conference on Learning Representations
- Litmaps
- NovGauge
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →