发表机构
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散语言模型的水印方法与迭代并行去掩蔽不兼容的问题,提出SAC-Copula方法,通过高斯Copula构建平滑相关Gumbel扰动场实现质量保留水印,在LLaDA等数据集上验证了其质量-可检测性权衡优势。
AI 中文摘要
对扩散语言模型(DLM)进行水印处理需要与迭代并行去掩蔽而非自回归解码兼容的机制。现有基于采样的水印方法通常注入逐位置独立同分布(i.i.d.)扰动,这可能与DLM解码动态 poorly 对齐并降低生成质量。我们提出SAC-Copula,一种基于高斯Copula构建的平滑、局部相关Gumbel扰动场的DLM质量保留水印方法。我们还开发了使用协方差感知滤波和原生样本校准的SAC感知检测器。机制层面分析表明,局部相关性降低了潜在扰动粗糙度,更好地匹配迭代细化动态。在LLaDA上的实验显示,与现有基线相比,SAC-Copula实现了有利的质量-可检测性权衡。特别是,对Dream-7B和额外数据集的进一步评估表明,SAC-Copula较i.i.d. Gumbula基线显著提高了PPL尾部稳定性,同时保持了强低FPR可检测性和有竞争力的整体生成质量。额外的token编辑压力测试进一步评估了水印在受控同步漂移下的鲁棒性。
英文摘要
Watermarking diffusion language models (DLMs) requires mechanisms compatible with iterative parallel unmasking rather than autoregressive decoding. Existing sampling-based watermarking methods typically inject position-wise i.i.d. perturbations, which can be poorly aligned with DLM decoding dynamics and degrade generation quality. We propose SAC-Copula, a quality-preserving watermarking method for DLMs based on smooth, locally correlated Gumbel perturbation fields constructed via a Gaussian copula. We further develop a SAC-aware detector using covariance-aware filtering and native-sample calibration. Mechanism-level analysis shows that local correlation reduces latent perturbation roughness and better matches iterative refinement dynamics. Experiments on LLaDA show that SAC-Copula achieves a favorable quality-detectability trade-off compared with existing baselines. In particular, further evaluations on Dream-7B and additional datasets show that SAC-Copula substantially improves PPL tail stability over the i.i.d. Gumbel baseline, while maintaining strong low-FPR detectability and competitive overall generation quality. Additional token-edit stress tests further assess watermark robustness under controlled synchronization drift. Code is available at https://github.com/PunkyKnife/SAC-Copula.
Comments24 pages, 14 figures. Accepted to Findings of EMNLP 2026