A new arXiv paper explores the internal mechanisms that cause Large Language Models (LLMs) to admit their mistakes. Researchers found that LLMs rarely retract incorrect answers, even when they can recognize the error in a separate interaction. A key predictor of retraction is the model's internal "belief" state during generation, which can diverge from its parametric knowledge. Experiments showed that influencing this belief state can causally drive retraction, and that supervised fine-tuning primarily improves retraction by enhancing the accuracy of these internal beliefs. AI
IMPACT This research highlights a critical limitation in current LLMs, suggesting that improving their ability to admit mistakes may require focusing on their internal 'belief' mechanisms rather than just factual knowledge.
RANK_REASON The cluster contains a single academic paper published on arXiv detailing research findings. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- LLMs
- ScienceCast
- Yuqing Yang
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →