A new research paper explores the discrepancy between what language models know internally and what they express through their outputs. The study found that while linear probes on internal model states can accurately predict a model's knowledge, the model's actual behavior often fails to reflect this knowledge due to miscalibrated decision thresholds. This issue, particularly a single scalar offset, can erase correct answers even when the model's internal representations are accurate. The research suggests that simple corrections, such as adjusting decision thresholds or using calibrated margin decoding, can significantly improve a model's behavioral accuracy without retraining. AI
IMPACT Highlights a critical gap in LLM interpretability, suggesting methods to improve model reliability and align internal knowledge with external behavior.
RANK_REASON Research paper detailing findings on language model behavior and internal knowledge. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →