Contrastive Activation Addition
PulseAugur coverage of Contrastive Activation Addition — every cluster mentioning Contrastive Activation Addition across labs, papers, and developer communities, ranked by signal.
-
New research questions effectiveness of activation steering in language models
A new research paper explores the phenomenon of activation steering in language models, questioning whether observed gains reflect intended control or compatibility with answer encodings. The study introduces Cross-Enco…
-
New method probes what activation steering truly controls in language models
Researchers have introduced a new evaluation method called Cross-Encoding Steering Evaluation to better understand what activation steering controls in language models. This method aims to distinguish between genuine co…
-
New method detects misinformation by analyzing LLM internal representations
Researchers have developed a novel method for detecting misinformation by analyzing the internal representations of language models, rather than relying on external knowledge or surface-level text features. This approac…
-
New CircuitSteer framework enhances LLM control via multi-layer semantic circuits
Researchers have developed a new framework called CircuitSteer, which uses Sparse Autoencoders to identify and manipulate specific semantic circuits within multiple layers of large language models. This method allows fo…