PulseAugur
EN
LIVE 18:59:27

LLMs Under-Confident in Recommendations, Study Finds

A new study auditing four large language models—Mistral Large, Llama 3.3 70B Instruct, GPT-OSS 120B, and Claude Sonnet 4.6—reveals that these models are systematically under-confident when asked to recommend items from specific catalogs. While previous research focused on over-confidence in hallucinations, this audit found that the models often express lower confidence than their actual accuracy would suggest, even when not hallucinating. The study suggests this under-confidence stems from a mismatch in how prompts elicit confidence ratings, rather than a true lack of certainty about catalog membership. AI

IMPACT Highlights a potential mismatch in how LLMs express confidence, suggesting current methods may not accurately reflect their understanding of catalog-specific recommendations.

RANK_REASON Academic paper detailing an audit of LLM confidence calibration.

Read on arXiv cs.IR (Information Retrieval) →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

LLMs Under-Confident in Recommendations, Study Finds

COVERAGE [2]

  1. arXiv cs.CL TIER_1 English(EN) · Srijith Ravikumar ·

    Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

    arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-$K$ item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate; none ask whether the model knew it was hallucinating. We jointly audit halluc…

  2. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Srijith Ravikumar ·

    Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

    LLM recommenders for top-$K$ item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate; none ask whether the model knew it was hallucinating. We jointly audit hallucination rate (OOD@10) and verbalized-confidence ca…