Researchers have developed a new method called StateSwap to investigate how large language models (LLMs) process multiple-choice questions inconsistently based on prompt framing. By introducing a special token, [STATE], and analyzing its activation in intermediate layers, they found that support-oriented and elimination-oriented framings induce distinct internal representations. Swapping these [STATE] activations between prompts can alter model predictions and improve agreement across different framings, suggesting these internal states are behaviorally relevant. AI
IMPACT Provides a new technique for understanding and potentially improving the consistency of LLM responses to varied prompts.
RANK_REASON The cluster describes a new research paper detailing a novel method for probing LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →