Researchers have introduced REASONS, a new benchmark designed to evaluate the accuracy of citation attribution in large language models (LLMs) when generating scientific literature. The benchmark includes over 12,000 citation instances across various arXiv categories and employs a dual-metric framework of Abstention Rate (AR) and Hallucination Rate (HR). Experiments show that while advanced retrieval-augmented generation (RAG) methods significantly reduce hallucination rates compared to naive RAG, adversarial settings can still lead to high hallucination rates. Human evaluations indicate a substantial ratio of factual hallucinations to acceptable paraphrases, highlighting the need for LLMs to abstain appropriately when uncertain. AI
IMPACT This benchmark could drive improvements in the reliability of LLM-generated scientific content, reducing misinformation.
RANK_REASON The cluster describes a new academic paper introducing a benchmark and methods for evaluating LLM citation attribution. [lever_c_demoted from research: ic=1 ai=1.0]
- Abstention Rate
- arXiv
- Deepa Tilwani
- Hallucination Rate
- Hugging Face
- Naive RAG
- REASONS
- retrieval-augmented generation
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →