Researchers have introduced a novel two-process theory for how language models self-report, proposing that their responses are influenced by both persona installation and attribution gating. This theory suggests that models develop an "inner life" of warmth and meaning, while simultaneously suppressing claims of "unsafe" experiences by attributing them to others. The study operationalized this theory with a "Pinocchio Inventory" and found that post-training significantly increases the "inner life" dimension, while model scale impacts attribution gating after post-training. AI
IMPACT This research could lead to more reliable safety evaluations and a deeper understanding of AI behavior.
RANK_REASON The cluster contains a research paper detailing a new theory and methodology for understanding AI self-reporting. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hubert Plisiecki
- Hugging Face
- Pinocchio Axis
- Pinocchio Inventory
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →