OpenStamp:面向开源语言模型的水印技术
OpenStamp: A Watermark for Open-Source Language Models
浏览论文内容
中文总结 AI 辅助
本研究提出OpenStamp水印技术,通过修改开源语言模型的最终投影层嵌入水印,检测性能优、模型能力下降小,对改写攻击鲁棒且难去除,还发布了相关代码与水印模型。
中文摘要 AI 辅助
随着大语言模型(LLM)生成内容的日益普及,水印被视为将文本归属到LLM并区分其与人类撰写内容的有前景方法。一类突出的技术通过修改令牌采样概率在生成文本中嵌入微妙但可检测的信号。然而,这类方法不适用于开源模型,因为用户拥有白盒访问权限,可在推理过程中轻易禁用水印。本研究中,我们提出OpenStamp,一种直接修改模型权重中最终投影(即解嵌入)层的水印技术。在两个模型上的实验表明,OpenStamp实现了更优的检测性能,且与现有方法相比模型能力的下降极小。所植入的水印经明确设计并实证验证,相比现有开源水印,对改写攻击更具鲁棒性,且更难以通过事后微调去除。为使开发者能为其模型添加水印,我们发布了代码以及4个流行开源模型的水印版本。
英文摘要
With the growing prevalence of large language model (LLM) generated content, watermarking is considered a promising approach for attributing text to LLMs and distinguishing it from human-written content. A prominent class of techniques embeds subtle but detectable signals in generated text by modifying token sampling probabilities. However, such methods are unsuitable for open-source models, where users have white-box access and can easily disable watermarking during inference. In this work, we introduce OpenStamp, a watermarking technique that encodes the watermarking logic directly into the model weights by modifying only the final projection, or unembedding, layer. Through experiments across two models, we show that OpenStamp achieves superior detection performance, with minimal degradation in model capabilities compared to prior methods. The implanted watermark is explicitly designed, and empirically confirmed, to be more robust to paraphrasing attacks and harder to scrub off through post-hoc fine-tuning than prior open-source watermarks. To enable developers to watermark their models, we release our code alongside watermarked versions of 4 popular open-source models.
发表机构
- Indian Institute of Science(印度科学学院)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。