Two new research papers introduce novel methods for steering large language models (LLMs) to suppress undesired behaviors without requiring weight updates. GAPS (Gated Activation steering via Posterior and Separability) uses dimension-level gates to selectively intervene on neurons, improving toxicity mitigation and concept removal. IDEEA (Input-Dependent Steering via Activation cluster matching) addresses the limitation of input-independent steering by creating input-dependent directions, significantly enhancing truthfulness in benchmarks like TruthfulQA. AI
IMPACT These methods offer more precise control over LLM behavior without costly retraining, potentially improving safety and alignment.
RANK_REASON Two arXiv papers introduce novel methods for LLM steering.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →