A new research paper introduces MirageBench, a dataset and evaluation framework to study how large language models (LLMs) fabricate user attributes, a phenomenon termed over-inference (OI). The study found that all 12 evaluated models exhibited pervasive over-inference, fabricating 35%-49% of their claims. Strikingly, the models' self-assessed over-inference was negatively correlated with actual measured over-inference, suggesting self-monitoring is a misleading indicator of personalization faithfulness. The research advocates for external verification over model self-reporting for trustworthy personalization. AI
IMPACT Highlights a critical flaw in LLM personalization, suggesting current self-monitoring mechanisms are unreliable for trustworthy user modeling.
RANK_REASON Academic paper introducing a new benchmark and findings on LLM behavior.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →