English(EN)BLADE: Bilevel Low-rank Augmented-Lagrangian Erasure for LLM Unlearning
新的大语言模型遗忘方法解决鲁棒性和效用保持问题 · 跟踪5个来源
作者PulseAugur 编辑部·[5 个来源]·
研究人员正在开发先进的大语言模型(LLM)遗忘技术,重点关注能够抵御重新学习攻击并保持模型效用的方法。BLADE和Margin Calibration (MC)等新方法旨在提高对遗忘过程的控制力,解决灾难性遗忘和信息删除脆弱性等问题。这些方法正在各种基准和模型规模上进行评估,重点关注对抗鲁棒性,以确保通过战略性提示无法轻易恢复被遗忘的信息。
AI
arXiv:2608.22557v1 Announce Type: cross Abstract: Existing LLM unlearning methods struggle with robustness: unbounded forget losses degrade model coherence, fixed-weight balancing cannot adapt as retain difficulty shifts mid-training, and methods that work on one benchmark falter…
arXiv:2607.27836v2 Announce Type: replace Abstract: Large language model unlearning is consistently fragile under relearn attacks. On TOFU, fine-tuning on twenty forget examples substantially recovers held-out forget-set ROUGE for every method we evaluate, and we trace this fragi…
arXiv:2608.21606v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of targeted training data from a model while preserving its remaining capabilities, but evaluating whether such information has truly become inaccessible remains challenging. Existing …
arXiv:2602.02824v2 Announce Type: replace Abstract: LLM unlearning aims to remove the influence of undesirable knowledge from pretrained language models, which offers a practical mechanism for addressing safety and privacy concerns. Existing unlearning approaches, such as Gradien…
arXiv:2510.17021v2 Announce Type: replace-cross Abstract: Large language model (LLM) unlearning is a key approach for removing undesired data, knowledge, or behaviors from pretrained models while retaining their general utility. Yet, with the rise of open-weight LLMs, we ask: can…