Researchers have introduced a new evaluation method called Cross-Encoding Steering Evaluation to better understand what activation steering controls in language models. This method aims to distinguish between genuine control over judgments and mere compatibility with answer identifiers. Experiments on NormBank revealed that a technique called Contrastive Activation Addition (CAA) primarily influences extraction indices rather than semantic labels, a phenomenon termed extraction-index following. This effect is more pronounced in later layers of the model and is also observed in Inference-Time Intervention methods, though results vary across different datasets like MNLI and Social Chemistry 101. AI
IMPACT Introduces a more rigorous evaluation method for understanding internal model mechanisms, potentially leading to more reliable control over LLM outputs.
RANK_REASON The item is a research paper detailing a new method for evaluating language model behavior. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
- activation steering
- Contrastive Activation Addition
- Cross-Encoding Steering Evaluation
- MNLI
- Social Chemistry 101
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →