A new research paper published on arXiv challenges the effectiveness of activation steering in language models. The study found that steering a model towards a specific behavior, such as politeness, does not isolate that behavior but instead pulls the model towards a default set of favored behaviors like refusal, sycophancy, and poeticism. This interference effect was observed across ten different instruction-tuned models and 24 distinct behaviors, with the pull towards defaults being stronger in smaller models. AI
IMPACT Challenges the modularity assumption in current language model control techniques, suggesting limitations in fine-grained behavioral steering.
RANK_REASON Research paper published on arXiv detailing findings about language model behavior. [lever_c_demoted from research: ic=1 ai=1.0]
- activation steering
- arXiv
- Hugging Face
- Instruction-Tuned Models
- language model
- poeticism
- refusal
- Steering Interference Reflects the Model's Defaults, Not the Behavior Directions
- sycophancy
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →