A new research paper explores the distinction between local and global confidence signals in autoregressive language models. The study found that these two measures, derived from the probability of the greedy-selected answer token versus the frequency of answers from repeated sampling, are weakly correlated and differ in their association with correctness. Global confidence showed a moderate link to accuracy, while local confidence had little correlation. The research also indicated that disagreements between these confidence signals can signal sampling instability in models, particularly on benchmarks like the ARC challenge. AI
IMPACT Highlights the need for careful interpretation of model confidence metrics, impacting how AI systems are evaluated and controlled.
RANK_REASON Research paper published on arXiv detailing findings about confidence signals in language models. [lever_c_demoted from research: ic=1 ai=1.0]
- ARC challenge
- autoregressive language models
- Julio Amador Diaz Lopez
- Massive Multitask Language Understanding
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →