Two research papers explore methods for improving the reliability of answers generated by large language models (LLMs), particularly in question-answering tasks. The first paper introduces A-CRC-QA, a post-hoc calibration framework designed to control error rates among accepted answers by reformulating selection-conditioned error control as a linear expectation constraint. The second paper empirically studies how to derive calibrated confidence estimates from token probabilities for mathematical question answering, comparing single-pass and multi-pass estimators and evaluating post-hoc calibration methods like Platt scaling and isotonic regression. AI
IMPACT Enhances the reliability of LLM outputs, crucial for applications requiring high accuracy and trustworthiness.
RANK_REASON Two academic papers published on arXiv presenting novel methods for LLM uncertainty quantification and calibration.
- arXiv
- Hugging Face
- Isotonic regression
- large language models
- Monte Carlo Dropout
- Platt scaling
- A-CRC-QA
- MedMCQA
- question answering
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →