A new study auditing four large language models—Mistral Large, Llama 3.3 70B Instruct, GPT-OSS 120B, and Claude Sonnet 4.6—reveals that these models are systematically under-confident when asked to recommend items from specific catalogs. While previous research focused on over-confidence in hallucinations, this audit found that the models often express lower confidence than their actual accuracy would suggest, even when not hallucinating. The study suggests this under-confidence stems from a mismatch in how prompts elicit confidence ratings, rather than a true lack of certainty about catalog membership. AI
IMPACT Highlights a potential mismatch in how LLMs express confidence, suggesting current methods may not accurately reflect their understanding of catalog-specific recommendations.
RANK_REASON Academic paper detailing an audit of LLM confidence calibration.
Read on arXiv cs.IR (Information Retrieval) →
- Amazon Reviews 2023 Toys
- Claude Sonnet 4.6
- GPT-OSS 120B
- Llama 3.3 70B Instruct
- Mistral Large
- MovieLens-25M
- Yelp Open Dataset
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →