Researchers have developed a method to detect when large language models (LLMs) are answering questions they cannot truly answer, or responding prematurely in a dialogue. They created a new benchmark and evaluation harness to test this capability across six datasets and six open-weight LLMs. The findings indicate that signals for unanswerability transfer well between similar datasets, such as those involving missing information in math problems or text passages, but transfer poorly to different types of unanswerability like epistemic "known-unknowns". While a calibrated probe can accurately identify underspecified turns without model fine-tuning, its end-task success is limited, suggesting the remaining gap lies in how models utilize clarification rather than detection. AI
IMPACT This research could lead to LLMs that are more reliable in identifying and refusing to answer unanswerable questions, improving dialogue systems and information retrieval.
RANK_REASON The cluster contains a research paper detailing a new method for evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- large language models
- Musique
- ScienceCast
- SQuAD2.0
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →