PulseAugur
实时 15:12:30
English(EN) DisarmRAG: Stealthy Retriever-Centric Poisoning to Disable Self-Correction in Retrieval-Augmented Generation (Extended Version)

新的DisarmRAG攻击禁用RAG系统中的LLM自我纠正

研究人员开发了一种名为DisarmRAG的新方法来利用检索增强生成(RAG)系统中的漏洞。与以往针对知识库的攻击不同,DisarmRAG会破坏检索器组件,注入指令来禁用大型语言模型的自我纠正能力。这种新颖的方法利用迭代协同优化过程和隐蔽的模型编辑技术来确保有效性,在多个LLM和基准测试中实现了超过90%的成功率,甚至能抵御检测防御。 AI

影响 突显了RAG系统中的新漏洞,可能影响LLM部署的可靠性和安全性。

排序理由 详细介绍LLM系统新型攻击方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的DisarmRAG攻击禁用RAG系统中的LLM自我纠正

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Yanbo Dai, Zhenlan Ji, Zongjie Li, Kuan Li, Shuai Wang ·

    DisarmRAG:一种隐蔽的以检索为中心的投毒方法,用于禁用检索增强生成中的自我纠正(扩展版)

    arXiv:2508.20083v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) has become a standard approach for improving the reliability of large language models (LLMs). Prior work demonstrates the vulnerability of RAG systems by misleading them into generating…