Two new research papers explore the complexities of large language model (LLM) refusals. The first paper, "A Unified Mechanistic Analysis of Knowledge- and Safety-Based Refusals," suggests that while knowledge-based and safety-based refusals share underlying mechanisms, they diverge in later layers, with safety refusals showing a stronger transfer to knowledge refusals. The second paper, "The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior," demonstrates that LLM safety evaluations are unreliable due to inconsistencies introduced by random seeds and temperature settings, with a significant percentage of prompts exhibiting decision flips. AI
IMPACT Highlights the need for more robust LLM safety evaluation protocols that account for stochastic variations in model behavior.
RANK_REASON Two academic papers published on arXiv analyzing LLM refusal mechanisms and stability.
- Claude 3.5 Haiku
- Gemma-3-12B
- Llama-3.1:8b
- LLaMA-70B
- Qwen 2.5 7B
- Qwen 3:8b
- arXiv
- knowledge-based refusal
- safety-based refusal
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →