arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Anchor-ECC:通过纠错码对带水印的大语言模型输出进行局部完整性校验

Anchor-ECC: Local Integrity Checking for Watermarked LLM Outputs via Error-Correcting Codes

Zewei Deng, Muhammad Siddeek, Liyan Xie, Mohamed Seif, Mengdi Wang, H. Vincent Poor, Andrea Goldsmith

arXiv 2609.38722首次发表:更新:

发表机构

University of Minnesota; Google; Oakland University; Princeton University; Stony Brook University(明尼苏达大学; 谷歌; 奥克兰大学; 普林斯顿大学; 石溪大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Anchor-ECC,通过纠错码约束和边界锚点结合动态规划解码器,实现对大语言模型水印输出的局部编辑检测与定位,在多个模型上达到高检测率与低误报率。

AI 中文摘要

大语言模型水印已成为一种有效方法,通过在生成过程中嵌入可检测的模式来区分AI生成的文本与人类撰写的文本。然而,生成后的小幅编辑可能改变文本含义而不移除其整体水印信号,造成修改后的内容仍被归因于原始模型的风险。我们提出了Anchor-ECC,该方法将纠错码(ECC)约束和显式边界锚点纳入水印结构,并将其与动态规划解码器配对,以检测和定位生成后的编辑。在Qwen3-8B、Mistral-7B-Instruct-v0.3和OPT-125M上,近似硬设置实现了约99.7%的块级真正例率(TPR),在混合插入、删除和替换下,编辑检测的误报率(FAR)最多为7.6%,同时保持了水印输出与未水印文本之间的区分度。额外的质量实验确定了具有较低困惑度的配置,这些配置在保持强编辑检测性能的同时,支持检测可靠性与生成质量之间的可配置权衡。综合来看,这些结果将大语言模型水印从来源识别扩展到局部完整性验证,并支持检测可靠性与生成质量之间的可配置权衡。

英文摘要

LLM watermarking has become an effective approach to distinguishing AI-generated text from human-written text by embedding detectable patterns during generation. However, a small post-generation edit may change the meaning of the text without removing its overall watermark signal, creating a risk that the modified content is still attributed to the original model. We propose Anchor-ECC, which incorporates the error-correcting code (ECC) constraints and explicit boundary anchors into the watermark structure and pairs them with a dynamic-programming decoder to detect and localize post-generation edits. Across Qwen3-8B, Mistral-7B-Instruct-v0.3, and OPT-125M, the approximate-hard setting achieves about 99.7% block-level true positive rate (TPR) with at most 7.6% false alarm rate (FAR) for edit detection under mixed insertions, deletions, and substitutions, while preserving the distinction between watermarked outputs and unwatermarked text. Additional quality experiments identify lower-perplexity configurations that retain strong edit-detection performance. Together, these results extend LLM watermarking from source identification to local integrity verification while supporting configurable trade-offs between detection reliability and generation quality.

Comments18 pages, including references and appendices; 1 figure and 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑