PulseAugur
EN
LIVE 13:43:07
ENTITY Arditi et al

Arditi et al

PulseAugur coverage of Arditi et al — every cluster mentioning Arditi et al across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
0
3 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
3 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 3 TOTAL
  1. TOOL · CL_91338 ·

    AI Refusal Control: DiM vs. INLP Methods Compared

    Researchers have compared two methods, Diff-in-Means (DiM) and Iterative Nullspace Projection (INLP), for controlling refusal behavior in AI chat models. The study found that INLP's counterfactual flipping intervention …

  2. TOOL · CL_90016 ·

    Sloppy AI Abliteration Costs More Than Technique Itself

    A recent analysis explores the cost of "abliteration," a technique to remove refusal capabilities from AI models. The author investigates whether the performance degradation observed in abliterated models is inherent to…

  3. RESEARCH · CL_50584 ·

    New research audits LLM alignment shifts using effective rank

    A new research paper introduces an "effective-rank" audit to analyze how alignment techniques alter the internal workings of large language models. The study examines three open-weight models: Llama-3.1-8B-Instruct, Gem…