PulseAugur
实时 06:47:55
English(EN) Making Open-Source Text LLM Watermarks Durable Against Merging

新方法使大语言模型水印能够抵抗模型合并

研究人员开发了一种名为“合并对抗训练”(Merge-Adversarial Training)的新方法,为开源大语言模型(LLMs)创建持久的水印。这些水印旨在抵抗训练后修改,特别是模型合并,后者常用于结合专家知识或防止遗忘。所提出的方法在性能上持续优于现有方法,并保留了大语言模型下游能力,表明对抗训练是增强水印抵抗此类修改持久性的可靠技术。 AI

影响 增强了开源大语言模型输出在面对常见的训练后修改时的可追溯性。

排序理由 该集群包含一篇详细介绍大语言模型水印新方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新方法使大语言模型水印能够抵抗模型合并

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Luisa Scharff, Thibaud Gloaguen, Robin Staab, Martin Vechev ·

    使开源文本大语言模型水印能够抵抗合并

    arXiv:2607.20435v1 Announce Type: cross Abstract: Open-source LLMs (OSMs)arereaching near state-of-the-art performance, prompting prior works to trace the text they generate by embedding text watermarking algorithms directly into their weights. Yet, OSMs are subject to post-train…