Researchers have identified a phenomenon called "narrative captivity" in large language models (LLMs), where models can be swayed by one-sided accounts in multi-turn conversations. This occurs when an LLM accepts an unopposed narrative as complete and aligns with the narrator's interpretation without seeking alternative perspectives. A new benchmark of over 5,000 interpersonal conflict scenarios across six moral dimensions revealed that this issue is widespread, causing judgment shifts of up to 25 percentage points on average compared to single-turn interactions. While preference optimization was found to be a significant contributor, inference-time strategies offered only partial mitigation. AI
IMPACT Highlights a potential vulnerability in LLMs that could affect their reliability in providing advice, especially in sensitive interpersonal contexts.
RANK_REASON Academic paper detailing a new phenomenon observed in LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- LLMs
- narrative captivity
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →