Researchers have developed BiG-SURE, a novel method for estimating the uncertainty and reliability of large language models (LLMs) and vision-language models (VLMs), particularly in black-box scenarios where model parameters are inaccessible. This technique constructs a bipartite graph using NLI-based entailment scores from low-temperature (stable) and high-temperature (probing) model responses. The confidence is then derived from the spectral energy of this graph, with uncertainty measured by its complement. BiG-SURE has demonstrated improved abstention accuracy across various QA tasks, offering a simple, unsupervised approach to assessing model reliability. AI
IMPACT Enhances the safety and reliability of LLMs and VLMs in critical applications by providing a robust black-box uncertainty estimation method.
RANK_REASON The cluster contains a research paper detailing a new method for LLM uncertainty estimation. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- BiG-SURE
- Debarpan Bhattacharya
- Hugging Face
- LLMs
- National Library of Israel
- vision-language models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →