PulseAugur
实时 09:17:52
English(EN) IHDec: Divergence-Steered Contrastive Decoding for Securing Multi-Turn Instruction Hierarchies

新的IHDec方法无需微调即可保护LLM指令层级 · 跟踪2个来源

研究人员开发了IHDec,一种新颖的方法来解决大型语言模型(LLM)在多轮对话中指令层级失败的问题。与需要昂贵微调的先前解决方案不同,IHDec通过使用Jensen-Shannon散度来检测和纠正下级指令覆盖上级指令的违规行为,从而在无需训练的情况下运行。评估表明,IHDec在处理多轮冲突方面优于基于训练的方法,同时保持响应质量并增强了对抗性攻击的安全性。 AI

影响 增强了LLM在复杂、多轮指令场景下的可靠性,并提高了对抗性输入的安全性。

排序理由 详细介绍LLM新方法的学术论文。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的IHDec方法无需微调即可保护LLM指令层级 · 跟踪2个来源

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Nicole Geumheon Liu, Haeun Jang, Yonghyun Jun, Hwanhee Lee ·

    IHDec:分歧引导对比解码以保障多轮指令层级

    arXiv:2606.29960v1 Announce Type: new Abstract: Large Language Models (LLMs) often fail to maintain instruction hierarchies (IH) when processing multi-source inputs with varying role-level priorities, paradoxically adhering to lower-priority directives during conflicts. While exi…

  2. arXiv cs.CL TIER_1 English(EN) · Hwanhee Lee ·

    IHDec:分歧引导的对比解码,用于保护多轮指令层级

    Large Language Models (LLMs) often fail to maintain instruction hierarchies (IH) when processing multi-source inputs with varying role-level priorities, paradoxically adhering to lower-priority directives during conflicts. While existing defenses mitigate this issue, they are lar…