发表机构
Nanjing University of Science and Technology; Anhui University of Technology; The Hong Kong Polytechnic University; University of Luxembourg; CSIRO(南京理工大学; 安徽工业大学; 香港理工大学; 卢森堡大学; 澳大利亚联邦科学与工业研究组织)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究生产水印导致大语言模型幻觉的问题,提出令牌级和注意力级两种干预措施,在保持流畅性的同时将事实错误减少约90%,强调事实性应成为水印评估的首要标准。
AI 中文摘要
文本水印有助于识别AI生成的内容,但其对事实可靠性的影响仍未得到充分探索。在本文中,我们研究了水印幻觉:即使上下文中存在所需证据且未加水印的模型能够正确回答,水印仍会诱发或放大事实错误。通过受控的检索增强生成设置,我们在相同上下文、查询和解码配置下比较未加水印和水印生成的输出,并量化其事实准确性的下降。在六种代表性水印方法(包括KGW、SWEET、DiPmark、GumbelSoft、Gumbel-Max和SynthID水印)中,我们一致观察到水印诱发的幻觉。水印输出可能保持流畅,同时引入事实错误。我们将此失败模式归因于两种机制:(1)当前步骤中由方法特定的重新加权或键控采样引起的令牌扰动,以及(2)前缀引起的注意力漂移,该漂移通过自回归解码累积,并削弱后续对事实上下文的注意力。基于此分析,我们提出了两种即插即用的干预措施,分别作用于令牌和注意力层面,可集成到现有水印方法中以提高事实性。在匹配的TPR为0.90、FPR为1%时,结合这两种干预措施相对于仅水印解码可将事实错误减少约90%,同时保持流畅性和相当的解码效率。总体而言,这项工作强调事实性应作为水印评估中与可检测性和鲁棒性并列的首要标准,并呼吁在事实关键应用中部署水印前进行仔细的事实性验证。
英文摘要
Text watermarking helps identify AI-generated content, but its effect on factual reliability remains underexplored. In this paper, we study watermarking hallucination: factual errors induced or amplified by watermarking even when the required evidence is present in the context and the unwatermarked model can answer correctly. Using a controlled retrieval-augmented generation setting, we compare unwatermarked and watermarked generations under the same context, query, and decoding configuration, and quantify their factual accuracy decrease. Across six representative watermarking methods, including KGW, SWEET, DiPmark, GumbelSoft, Gumbel-Max, and SynthID watermarking, we consistently observe watermark-induced hallucination. Watermarked outputs can remain fluent while introducing factual errors. We attribute this failure mode to two mechanisms: (1) token perturbations in the current-step arising from method-specific reweighting or keyed sampling, and (2) prefix-induced attention drift, which accumulates through autoregressive decoding and weakens later attention to the factual context. Motivated by this analysis, we propose two plug-in interventions at the token and attention levels that can be integrated into existing watermarking methods to improve factuality. At a matched TPR of 0.90 at 1% FPR, combining the two interventions reduces factual errors by approximately 90% relative to watermark-only decoding while preserving fluency and comparable decoding efficiency. Overall, this work highlights factuality as a first-class criterion in watermark evaluation, alongside detectability and robustness, and calls for careful factuality validation before deploying watermarks in fact-critical applications.
CommentsAccepted at NeurIPS 2026