A new benchmark called SciFactCheck has been developed to evaluate the factuality of large language models (LLMs) when discussing scientific topics. The benchmark, which covers five scientific domains and identifies three types of hallucinations (unverifiability, overclaim, and attribution), found that models specifically fine-tuned on scientific data performed worse in factual reliability compared to their general-purpose counterparts. Furthermore, current fact-checking tools showed only moderate agreement with expert judgments on scientific content, highlighting a need for improved verification infrastructure. AI
IMPACT Challenges current domain-specific fine-tuning methods for LLMs and highlights the need for better scientific content verification infrastructure.
RANK_REASON The cluster contains a research paper detailing a new benchmark and evaluation of LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- LLMs
- Raia Abu Ahmad
- ScienceCast
- SciFactCheck
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →