arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37218cs.CRcs.AI

超越语义窄化:基于汉明邻域的鲁棒且高效的大语言模型水印

Beyond Semantic Narrowing: Robust and Efficient LLM Watermarking with Hamming Neighborhoods

Zewen Sun, Tongyang Zhao, Liyao Xiang, Mingxuan Ma, Lingzhe Wang, Zhiyuan Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有语义水印导致的语义窄化问题,提出HammingMark,利用前句语义哈希的汉明邻域定义水印有效性,保留更多自然延续,在C4和BookSum上实现强鲁棒性、高检测率,采样成本降低72.8%。

中文摘要 AI 辅助

语义水印通过将可检测信号嵌入句子级表示来提高对水印去除攻击的鲁棒性。然而,现有水印方法通常对生成的句子施加水印特定的语义偏好,而没有明确考虑大语言模型生成的高度非均匀且上下文相关的语义偏好。当这两种偏好对齐不佳时,许多自然延续变得与水印不兼容,导致语义窄化:语义自由度降低、重采样成本增加,以及在具有严格语义要求的任务上可能性能下降。为解决此问题,我们提出HammingMark,它使用前一句的语义哈希作为动态中心,并接受哈希落在其汉明邻域内的候选。在紧凑哈希空间中定义水印有效性为汉明邻域,保留了更大比例的自然可能语义延续。粗粒度的多对一哈希映射进一步允许多样化的语义实现保持水印有效。在C4和BookSum上的实验表明,HammingMark实现了强鲁棒性、高可检测性和接近无水印的生成质量,每个被接受的句子仅需2.2个采样候选,与现有采样效率最高的方法相比减少了72.8%。在具有严格语义约束的更复杂任务上,HammingMark实现了最高的检测率,并取得最高或并列最高的ROUGE-L分数,证明了其在受限生成设置下平衡水印可检测性和生成质量的有效性。

英文摘要

Semantic watermarking improves robustness against watermark removal attacks by embedding detectable signals into sentence-level representations. However, existing watermarking methods typically impose watermark-specific semantic preferences on generated sentences without explicitly accounting for the highly non-uniform and context-dependent semantic preference of LLM generation. When these two preferences are poorly aligned, many natural continuations become incompatible with the watermark, causing semantic narrowing: reduced semantic freedom, increased resampling cost, and potential degradation on tasks with strict semantic requirements. To alleviate this problem, we propose HammingMark, which uses the semantic hash of the preceding sentence as a dynamic center and accepts candidates whose hashes fall within its Hamming neighborhood. Defining watermark validity over a Hamming neighborhood in compact hash space retains a larger fraction of naturally likely semantic continuations. The coarse many-to-one hash mapping further allows diverse semantic realizations to remain watermark-valid. Experiments on C4 and BookSum show that HammingMark achieves strong robustness, high detectability, and near-unwatermarked generation quality, requiring only 2.2 sampled candidates per accepted sentence,a 72.8% reduction compared with the most sampling-efficient existing method. On more complex tasks with strict semantic constraints, HammingMark achieves the highest detection rates with the highest or tied-highest ROUGE-L scores, demonstrating its effectiveness in balancing watermark detectability and generation quality under constrained generation settings.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • Shanghai Innovation Institute(上海创新研究院)
  • Northwest Polytechnical University(西北工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑