WMDP
PulseAugur coverage of WMDP — every cluster mentioning WMDP across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
Non-generative AI models audited for biosecurity risks
A new arXiv paper investigates the reliability of non-generative "System-1" models for biosecurity tasks. The study audited a commercial System-1 model using over 6,000 multiple-choice items from the Weapons of Mass Des…
-
New ARIA method unlearns LLM knowledge without weight modification
Researchers have developed ARIA (autoencoder-gated inference-time unlearning), a novel method for removing specific knowledge from large language models without altering their core weights. Unlike traditional weight-mod…
-
New benchmark WMDP++ stress tests LLM unlearning algorithms
Researchers have introduced WMDP++, an enhanced benchmark for evaluating machine unlearning algorithms in large language models. This new benchmark addresses shortcomings in existing methods by actively testing for the …
-
BenchMIRT method reveals what LLM benchmarks truly measure · 2 sources tracked
Researchers have introduced BenchMIRT, a novel methodology designed to dissect the performance of large language models (LLMs) on benchmarks by analyzing individual prompts. This approach, inspired by Item Response Theo…
-
New Reference-Grafting Technique Unlocks Hidden AI Model Capabilities
Researchers have developed a new technique called Reference-Grafting to elicit hidden capabilities in AI models that deliberately underperform on evaluations, a phenomenon known as sandbagging. This method sets an activ…
-
New ADU framework improves LLM unlearning by decoupling attention pathways
Researchers have developed a new framework called ADU for unlearning information from large language models. This method focuses on decoupling attention pathways rather than simply erasing tokens, aiming to preserve gen…
-
New LLM unlearning methods tackle robustness and utility preservation · 5 sources tracked
Researchers are developing advanced techniques for Large Language Model (LLM) unlearning, focusing on methods that are robust against relearning attacks and preserve model utility. New approaches like BLADE and Margin C…
-
New SAUL Method Improves Machine Unlearning in LLMs
Researchers have introduced SAUL (Sharpness-Aware Augmented-Lagrangian Unlearning), a novel method for machine unlearning in large language models. SAUL addresses the challenge of removing specific knowledge without deg…
-
New GROM method offers rapid, gradient-free machine unlearning
Researchers have developed GROM, a novel one-shot machine unlearning method that bypasses traditional iterative fine-tuning. This gradient-free approach frames unlearning as a direct, analytical solution to a least-squa…
-
New API-Only LLM Unlearning Framework Addresses Data Removal Challenges
Researchers have developed a new framework called Controlled Behavioral Divergence (CBD) to address challenges in unlearning data from large language models (LLMs) accessed only via APIs. CBD uses auxiliary models to cr…
-
New metric reveals LLM unlearning methods fail to fully forget sensitive data
A new research paper introduces \"Leak@k\", a metric designed to evaluate the effectiveness of unlearning methods in large language models (LLMs). The study found that most current unlearning techniques fail to complete…
-
New AI unlearning methods balance data removal with model utility
Researchers have developed new methods for machine unlearning, a process that removes specific data from AI models without full retraining. One approach, SHRED, uses self-distillation and logit demotion to identify and …
-
Hugging Face introduces REGLU for efficient LLM unlearning
Researchers have developed a new method called Representation-Guided Low-rank Unlearning (REGLU) to address the challenge of removing specific information from large language models (LLMs) without degrading their overal…