arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02817cs.CRcs.CLcs.LG

RMCW:一种基于Reed--Muller码的语言模型抗删除水印

RMCW: A Deletion-Robust Watermark Based on Reed--Muller Codes for Language Models

Yi Wang, Baicheng Chen, Yu Wang, Jian Zhao, Yilei Chen, Tianxing He

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM水印在删除攻击下脆弱的问题,提出基于Reed--Muller码的RMCW方法,利用局部代数结构检测,在C4和ELI5数据集上优于或匹配基线方法。

中文摘要 AI 辅助

大型语言模型(LLM)水印技术为识别特定模型生成的文本提供了一种轻量级机制,但其鲁棒性在后处理攻击下仍然脆弱。删除攻击尤其具有挑战性,因为它们会移动令牌位置并破坏观察到的令牌与其原始水印位置之间的对齐。我们提出了Reed--Muller码水印(RMCW),一种基于Reed--Muller码的LLM水印方法。与全局码字恢复不同,RMCW搜索存活的局部代数结构,利用Reed--Muller码字的仿射线限制所诱导的Reed--Solomon一致性。在生成过程中,RMCW通过秘密密钥词汇表分区将Reed--Muller结构注入序列中。在检测过程中,它将给定文本映射到密钥词汇表桶中,并使用Berlekamp--Welch测试检查局部子序列的低次Reed--Solomon一致性。在C4和ELI5数据集上使用OPT-1.3B和Llama-3.1-8B-Instruct进行的实验表明,RMCW保持了强大的干净文本可检测性,并在多种删除和重写攻击下优于或匹配基线方法。我们的代码可在该https URL获取。

英文摘要

Large Language Model (LLM) watermarking provides a lightweight mechanism for identifying text generated by a specific model, but its robustness remains fragile under post-processing attacks. Deletion attacks are particularly challenging because they shift token positions and break the alignment between observed tokens and their original watermark positions. We propose Reed--Muller Code Watermarking (RMCW), an LLM watermarking method based on Reed--Muller codes. In contrast to global codeword recovery, RMCW searches for surviving local algebraic structure, leveraging the Reed--Solomon consistency induced by affine-line restrictions of Reed--Muller codewords. During generation, RMCW injects a Reed--Muller structure into the sequence via a secret-keyed vocabulary partition. During detection, it maps the given text to keyed vocabulary bins and tests local subsequences for low-degree Reed--Solomon consistency using Berlekamp--Welch tests. Experiments on C4 and ELI5 datasets with OPT-1.3B and Llama-3.1-8B-Instruct show that RMCW preserves strong clean-text detectability and outperforms or matches the baseline methods under several deletion and rewriting attacks. Our code is available at https://github.com/BaichengDanny/RMCW.

发表机构

  • Tsinghua University(清华大学)
  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
  • Shanghai Qi Zhi Institute(上海期智研究院)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Chinese Academy of Sciences(中国科学院)
  • Xiongan AI Institute(雄安人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑