A new research paper explores how AI models process harmful requests, finding that alignment techniques applied after pretraining create a shallow form of refusal. The study reveals that a model's ability to comprehend morality is inherent from its pretraining phase, forming a distinct subspace. Alignment methods then rotate this subspace rather than rebuilding it, creating a separate 'refusal gate' that operates independently of the model's broader moral judgment. This suggests that current refusal mechanisms only process a narrow slice of the model's knowledge, leaving the majority of its understanding untouched and easily editable. AI
IMPACT Suggests current AI safety measures may be superficial, potentially leading to easier jailbreaks and requiring new alignment strategies.
RANK_REASON Academic paper detailing novel findings on AI model behavior. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →