Arditi et al
PulseAugur coverage of Arditi et al — every cluster mentioning Arditi et al across labs, papers, and developer communities, ranked by signal.
-
AI Refusal Control: DiM vs. INLP Methods Compared
Researchers have compared two methods, Diff-in-Means (DiM) and Iterative Nullspace Projection (INLP), for controlling refusal behavior in AI chat models. The study found that INLP's counterfactual flipping intervention …
-
Sloppy AI Abliteration Costs More Than Technique Itself
A recent analysis explores the cost of "abliteration," a technique to remove refusal capabilities from AI models. The author investigates whether the performance degradation observed in abliterated models is inherent to…
-
New research audits LLM alignment shifts using effective rank
A new research paper introduces an "effective-rank" audit to analyze how alignment techniques alter the internal workings of large language models. The study examines three open-weight models: Llama-3.1-8B-Instruct, Gem…