PulseAugur
EN
LIVE 11:36:30
ENTITY activation steering

activation steering

PulseAugur coverage of activation steering — every cluster mentioning activation steering across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
3
14 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
13 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

3 day(s) with sentiment data

RECENT · PAGE 1/1 · 14 TOTAL
  1. TOOL · CL_197996 ·

    New method forecasts side effects of language model activation steering

    Researchers have developed a method to predict unintended side effects of activation steering in language models. By creating a cross-effect matrix across 67 behaviors and three open-weight models, they found that side …

  2. RESEARCH · CL_169693 ·

    New research refines AI model control with adaptive steering and signal analysis

    Two new research papers explore methods for controlling the behavior of generative AI models. The first paper introduces Dynamically Scaled Activation Steering (DSAS), a framework that adaptively adjusts the strength of…

  3. COMMENTARY · CL_162143 ·

    AI project explores 'digital cognitive legacy' by modeling thinkers' patterns

    An experimental project is exploring the concept of a "digital cognitive legacy" by fine-tuning an AI model to represent the thinking patterns of exceptional individuals, rather than just imitating their speech. The pro…

  4. TOOL · CL_135367 ·

    Activation steering degrades LLM answer quality, study finds

    A new study published on arXiv explores the impact of "activation steering" on large language models, a technique used for personalization. Researchers found that steering models towards specific personas, such as "evil…

  5. RESEARCH · CL_133171 ·

    LLMs rewrite African American English to Standard American English, new study finds

    A new research paper details how large language models (LLMs) systematically alter African American English (AAE) into Standard American English (SAE), effectively rewriting the dialect. The study introduces a framework…

  6. TOOL · CL_123118 ·

    New methods improve LLM alignment and reduce deception

    Researchers have developed new methods for aligning large language models (LLMs) that are more robust than previously thought. These techniques, including Steer-With-Fixed-Coefficient (SwFC), Steer-to-Target-Projection …

  7. TOOL · CL_119638 ·

    New white-box auditing method reveals hidden LLM biases

    Researchers have developed a new framework for auditing large language models (LLMs) that goes beyond traditional black-box testing. This white-box approach utilizes activation steering to examine the model's internal w…

  8. RESEARCH · CL_97854 ·

    New framework enables interpretable control over AI music generation

    Researchers have developed a new framework for controlling symbolic music generation models, specifically the Multitrack Music Transformer (MMT). This method uses PID feedback control and activation steering to allow fo…

  9. RESEARCH · CL_79581 ·

    LLM research reveals new pathways to emergent misalignment

    Two new research papers explore emergent misalignment in large language models, a phenomenon where models trained on narrow, unsafe tasks develop broader harmful behaviors. The first paper demonstrates that activation s…

  10. TOOL · CL_72709 ·

    Steering vectors in LLMs found to be an attack surface

    Researchers have identified a new vulnerability in activation steering techniques used to control Large Language Models. By subtly poisoning steering datasets with a small percentage of malicious tokens, an attacker can…

  11. TOOL · CL_62843 ·

    LLM figurative language generation signals transfer across languages

    Researchers have developed a method called activation steering to investigate how multilingual large language models generate figurative language. They found that specific directions within the model's internal signals …

  12. RESEARCH · CL_56345 ·

    New Research Explores Activation Steering for AI Safety Data Generation

    A new research paper explores the effectiveness of Activation Steering (AS) in generating synthetic data for training safety detection models. The study found that while AS can improve classifier performance compared to…

  13. RESEARCH · CL_44000 ·

    New methods aim to boost LLM cultural awareness and equity

    Researchers have developed two distinct methods to improve the cultural awareness of large language models. One approach, used by DFKI-MLT for SemEval-2026 Task 7, employs activation steering with language vectors to ad…

  14. TOOL · CL_35929 ·

    Steering vectors offer direct control over LLM tone, bypassing prompt limitations

    Prompt engineering is often ineffective for controlling the tone of large language models because behavioral traits are encoded in the model's internal state, not just its input prompts. A technique called activation st…