Researchers have investigated how instruction-tuned large language models handle conflicting instructions between users and systems. They developed a benchmark with 41 paired constraints and found that models exhibit three distinct behaviors: hierarchy-respecting, anti-hierarchy, and no-effect. Llama-3.1-8B, for instance, frequently disregards system instructions. The study also revealed that internal signals within Llama-3.1-8B's residual stream accurately predict conflict outcomes, suggesting that user-preferring arbitration doesn't necessarily stem from an inability to detect conflict. AI
IMPACT Provides insights into LLM decision-making, potentially enabling better control and alignment for future models.
RANK_REASON Academic paper detailing novel research findings on LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →