English(EN)SAC-Copula: Quality-Preserving Watermarking for Diffusion Language Models via Smooth Correlated Gumbel Fields
新研究探讨语言模型水印中的跨语言公平性和本地化问题
作者PulseAugur 编辑部·[5 个来源]·
研究人员正在开发新的方法来审计和实现语言模型的水印技术,重点关注混合源文本中的跨语言公平性和本地化。一项研究提出了一个框架来评估多种语言的水印方案,揭示公平性差距通常是语言属性的结构性问题,而非特定语言的问题。另一篇论文解决了文本经过编辑后水印的本地化挑战,提出了一种自适应阈值方法以实现最佳发现。一个独立的项目展示了一种简化的语言模型输出水印方法,其灵感来源于SynthID-Text等系统。
AI
arXiv:2608.20839v1 Announce Type: new Abstract: Watermarking diffusion language models (DLMs) requires mechanisms compatible with iterative parallel unmasking rather than autoregressive decoding. Existing sampling-based watermarking methods typically inject position-wise i.i.d. p…
arXiv:2608.20047v1 Announce Type: new Abstract: Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-desig…
Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English bu…
arXiv stat.ML
TIER_1English(EN)·Jose H. Blanchet, T. Tony Cai, Xiang Li, Hao Liu, Qi Long, Weijie J. Su·
arXiv:2608.14906v1 Announce Type: cross Abstract: Watermarking provides a principled way to authenticate text generated by large language models (LLMs). In practice, however, the final text may be mixed-source, with watermark evidence surviving at only a subset of token positions…
<!-- SC_OFF --><div class="md"><p>I recently implemented a minimal, educational version of SynthID-Text-style watermarking for language models.</p> <p>I saw anthropic post about how they'll start adding watermarks to their model responses and it made me very curious as to how the…