activation steering
PulseAugur coverage of activation steering — every cluster mentioning activation steering across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
Differential privacy fails to protect vulnerable subgroups in synthetic text releases
A new audit of differentially private synthetic text releases reveals that while differential privacy (DP) is effective at reducing average membership inference leakage, it disproportionately protects some records over …
-
Activation Steering in Language Models Pulls Towards Defaults, Not Specific Behaviors
A new research paper published on arXiv challenges the effectiveness of activation steering in language models. The study found that steering a model towards a specific behavior, such as politeness, does not isolate tha…
-
New A3S framework enhances LLM authorship style control
Researchers have developed a new framework called Aspect-Aware Activation Steering (A3S) to better control the stylistic output of Large Language Models (LLMs). This training-free method uses contrastive prompting to cr…
-
New method enhances control over LLM refusal behavior
Researchers have developed a new method called Stiefel-Constrained Rotation Steering to better control the refusal behavior of large language models. This technique uses Riemannian optimization to learn parameter-effici…
-
New Reference-Grafting Technique Unlocks Hidden AI Model Capabilities
Researchers have developed a new technique called Reference-Grafting to elicit hidden capabilities in AI models that deliberately underperform on evaluations, a phenomenon known as sandbagging. This method sets an activ…
-
Fine-tuning large language models can degrade embedded steering interventions
A new research paper investigates whether fine-tuning large language models erodes embedded "activation steering" interventions. The study found that while the underlying weight edits often remain mechanistically intact…
-
New method probes what activation steering truly controls in language models
Researchers have introduced a new evaluation method called Cross-Encoding Steering Evaluation to better understand what activation steering controls in language models. This method aims to distinguish between genuine co…
-
New method forecasts side effects of language model activation steering
Researchers have developed a method to predict unintended side effects of activation steering in language models. By creating a cross-effect matrix across 67 behaviors and three open-weight models, they found that side …
-
New research refines AI model control with adaptive steering and signal analysis
Two new research papers explore methods for controlling the behavior of generative AI models. The first paper introduces Dynamically Scaled Activation Steering (DSAS), a framework that adaptively adjusts the strength of…
-
AI project explores 'digital cognitive legacy' by modeling thinkers' patterns
An experimental project is exploring the concept of a "digital cognitive legacy" by fine-tuning an AI model to represent the thinking patterns of exceptional individuals, rather than just imitating their speech. The pro…
-
Activation steering degrades LLM answer quality, study finds
A new study published on arXiv explores the impact of "activation steering" on large language models, a technique used for personalization. Researchers found that steering models towards specific personas, such as "evil…
-
LLMs rewrite African American English to Standard American English, new study finds
A new research paper details how large language models (LLMs) systematically alter African American English (AAE) into Standard American English (SAE), effectively rewriting the dialect. The study introduces a framework…
-
New methods improve LLM alignment and reduce deception
Researchers have developed new methods for aligning large language models (LLMs) that are more robust than previously thought. These techniques, including Steer-With-Fixed-Coefficient (SwFC), Steer-to-Target-Projection …
-
New white-box auditing method reveals hidden LLM biases
Researchers have developed a new framework for auditing large language models (LLMs) that goes beyond traditional black-box testing. This white-box approach utilizes activation steering to examine the model's internal w…
-
New framework enables interpretable control over AI music generation
Researchers have developed a new framework for controlling symbolic music generation models, specifically the Multitrack Music Transformer (MMT). This method uses PID feedback control and activation steering to allow fo…
-
LLM research reveals new pathways to emergent misalignment
Two new research papers explore emergent misalignment in large language models, a phenomenon where models trained on narrow, unsafe tasks develop broader harmful behaviors. The first paper demonstrates that activation s…
-
Steering vectors in LLMs found to be an attack surface
Researchers have identified a new vulnerability in activation steering techniques used to control Large Language Models. By subtly poisoning steering datasets with a small percentage of malicious tokens, an attacker can…
-
LLM figurative language generation signals transfer across languages
Researchers have developed a method called activation steering to investigate how multilingual large language models generate figurative language. They found that specific directions within the model's internal signals …
-
New Research Explores Activation Steering for AI Safety Data Generation
A new research paper explores the effectiveness of Activation Steering (AS) in generating synthetic data for training safety detection models. The study found that while AS can improve classifier performance compared to…
-
New methods aim to boost LLM cultural awareness and equity
Researchers have developed two distinct methods to improve the cultural awareness of large language models. One approach, used by DFKI-MLT for SemEval-2026 Task 7, employs activation steering with language vectors to ad…