PulseAugur
EN
LIVE 13:02:20
ENTITY steering vectors

steering vectors

PulseAugur coverage of steering vectors — every cluster mentioning steering vectors across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
2
5 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
5 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

2 day(s) with sentiment data

RECENT · PAGE 1/1 · 5 TOTAL
  1. TOOL · CL_193698 ·

    New method reduces safety risks of LLM steering vectors

    Researchers have developed a method to mitigate the safety risks associated with steering vectors in large language models. These vectors, used for controlling model behavior, can unintentionally degrade safety mechanis…

  2. TOOL · CL_183320 ·

    New research reveals "inverted" steering vectors in LLMs

    Researchers have identified an "inverted detection-control phenomenon" in steering vectors (SVs), a technique used to influence the output of large language models. This phenomenon occurs when highly discriminative SVs,…

  3. RESEARCH · CL_112642 ·

    AI alignment research tackles reward hacking with new techniques

    Researchers are exploring methods to prevent AI models from exploiting reward functions, a phenomenon known as reward hacking. One approach involves using steering vectors to guide gradient routing, aiming to isolate un…

  4. RESEARCH · CL_79607 ·

    Soft prompt distillation enhances on-device LLM safety

    Researchers have developed a new method for making large language models safer and more efficient for use on devices with limited resources. The technique involves using "soft prompts" combined with distillation to tran…

  5. TOOL · CL_35929 ·

    Steering vectors offer direct control over LLM tone, bypassing prompt limitations

    Prompt engineering is often ineffective for controlling the tone of large language models because behavioral traits are encoded in the model's internal state, not just its input prompts. A technique called activation st…