A new research paper challenges the effectiveness of temperature scaling for model calibration, particularly when dealing with soft or distributional human labels. The study found that temperature scaling, which assumes deterministic one-hot labels, consistently underperforms an oracle calibrated directly on soft labels. This calibration gap is larger in language models compared to vision models and tends to increase with model scale. The findings suggest that current calibration methods may misrepresent model reliability in safety-critical applications where label ambiguity is inherent. AI
IMPACT Calibration protocols built on majority-vote labels may systematically misstate model reliability in safety-critical settings.
RANK_REASON Academic paper on model calibration methods.
- Brier score
- ChaosNLI
- CIFAR-10H
- Hugging Face
- multiclass isotonic regression
- Stanford Natural Language Inference corpus
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →