English(EN)Geometric Configurations of Perturbed Jailbreak Prompts
新的框架出现以评估和防御LLM越狱 · 跟踪4个来源
作者PulseAugur 编辑部·[5 个来源]·
研究人员正在开发新的方法来评估和防御大型语言模型(LLM)的越狱攻击。一种方法,不完整的提示越狱(IPJ),侧重于LLM如何延迟拒绝有害提示直到句子终止,并提出神经元级别的干预措施进行防御。另一个框架JailMeter使用信息瓶颈理论来更可靠地评估越狱的有效性,并取得了高精度。此外,Jailbreak Foundry提供了一个系统,将越狱论文翻译成可执行模块,用于可复现的基准测试,从而标准化不同模型和攻击的评估。
AI
arXiv:2607.20473v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests. Nevertheless, sentence completion remains vulnerable to incomplete harmful prompts. In this work, we formalize th…
arXiv cs.AI
TIER_1English(EN)·Lynn Delcon, Andres Algaba, Vincent Ginis·
arXiv:2607.20581v1 Announce Type: cross Abstract: Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this paper, we investigate the internal representations of suc…
arXiv:2607.19424v1 Announce Type: cross Abstract: The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates. We propose JailMeter, an evidence-base…
arXiv:2602.24009v4 Announce Type: replace-cross Abstract: Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale and difficult to compare across papers due to drift in datasets, harnesses, and judging protocols. We i…