arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12713cs.CRcs.AIcs.CL

追踪来源并检测互补大语言模型水印的篡改

Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks

  • Griffith University(格里菲斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiaoyan Feng, Yanjun Zhang, He Zhang, Leo Yu Zhang, Shirui Pan

AI总结:

本研究提出一种互补LLM水印方法,通过嵌入鲁棒与脆弱双信号,实现来源追踪与篡改检测,在两类LLM和两类提示数据集上,其篡改检测率优于现有方法,且保持了良好的归属鲁棒性与困惑度。

AI中文摘要:

对大语言模型(LLM)生成的文本加水印是追踪其来源的重要任务。现有LLM水印在编辑后仍能保留来源,但这种鲁棒性也让攻击者可篡改关键内容同时保留归属,此漏洞被称为“附随欺骗(piggyback spoofing)”。我们提出一种创新水印,可同时提供来源与篡改证据:它将鲁棒信号与脆弱信号共同嵌入每个生成的token,两种信号采用相同机制但使用独立密钥、在归一化文本上采用不同种子窗口,使一种信号抗编辑,另一种对读者可见的变化敏感。多轮无偏锦标赛重加权保留预期生成分布,而周期性轮次分配模式控制两种信号间的权衡。检测时,二者的分数构成二维空间,支持三种决策:完整(Intact)、已篡改(Tampered)、无水印(No-Watermark)。在两个大语言模型和两个提示数据集上,我们的方法在评估方法中展现出最高的篡改检测率,同时保持有竞争力的归属鲁棒性和困惑度(perplexity)。消融研究表明,可靠的三状态检测需要明确定义的完整性概念、两种信号的共同嵌入以及对编辑的互补敏感性。

英文摘要:

Watermarking LLM-generated text is an important task for tracing its provenance. Existing LLM watermarks preserve provenance under editing, but this same robustness allows an adversary to alter critical content while retaining attribution, a vulnerability known as piggyback spoofing. We introduce an innovative watermark that jointly provides provenance and tamper evidence. It co-embeds a robust signal and a fragile signal into each generated token. The signals share the same mechanism but use independent keys and different seeding windows over normalized text, making one resilient to edits and the other sensitive to reader-visible changes. Multiple rounds of unbiased tournament reweighting preserve the expected generation distribution, while a periodic round-allocation pattern controls the trade-off between the two signals. At detection, their scores form a two-dimensional space supporting three decisions: Intact, Tampered, and No-Watermark. Across two large language models and two prompt datasets, our method demonstrates the highest tamper-detection rate among the evaluated methods while maintaining competitive attribution robustness and perplexity. Ablation studies show that reliable three-state detection requires a well-defined notion of intactness, co-embedding of the two signals, and complementary sensitivity to edits.

补充信息

↑