EAP-IG
PulseAugur coverage of EAP-IG — every cluster mentioning EAP-IG across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Mechanistic Interpretability Research Reveals Objective-Level Recovery Gaps
A new research paper titled "Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability" has been published on arXiv. The paper investigates the effectiveness of current methods in mecha…
-
New C-ΔΘ method embeds LLM safety refusals into model weights
Researchers have developed a new method called C-ΔΘ (Circuit-Restricted Weight Arithmetic) to improve the safety of large language models. This technique aims to embed refusal capabilities directly into the model's weig…
-
New LLM Circuit Discovery Method Addresses Variances
A new research paper published on arXiv explores the variability in circuit discovery methods for Large Language Models (LLMs). The study identifies three main sources of variance: resampling, rephrasing, and sample-wis…