Two new research papers delve into the mechanisms behind making large language models refuse harmful requests. The first paper compares different post-training methods like supervised fine-tuning, reasoning-augmented fine-tuning, and ORPO across various models, finding that training methods significantly alter internal refusal computations. The second paper investigates representation steering, revealing that steering vectors primarily interact with the attention mechanism's OV circuit and largely ignore the QK circuit, with potential for significant sparsification without performance loss. AI
IMPACT Provides deeper understanding of LLM safety mechanisms, potentially leading to more robust and steerable AI systems.
RANK_REASON Two academic papers published on arXiv detailing research into LLM safety mechanisms.
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- IArxiv
- large language models
- ScienceCast
- Stephen Cheng
- Gemma 2 9B
- Llama-3.1:8b
- Optimized Ranking With Pairwise Observations
- Qwen3_8B
- reasoning-augmented fine-tuning
- supervised fine-tuning
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →