发表机构
Radboud University; University of Bristol; Faculty of Electrical Engineering and Computing, University of Zagreb(拉德堡德大学; 布里斯托大学; 萨格勒布大学电气工程和计算学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出利用MoE模型结构的WoE水印技术,可在黑盒场景下对被盗模型生成的文本进行来源归因,在1%假阳性率下平均真阳性率达90.1%,且具备较强的抗攻击能力。
AI 中文摘要
大语言模型(LLM)水印为文本来源追踪提供了一种机制,使模型所有者能够识别机器生成的内容并将其归因于特定的水印模型。然而,当前的LLM水印方法主要依赖推理时的采样器方法,且分析重点集中在稠密模型上。推理时方法仅在文本通过模型所有者控制的API明确生成时才有效,在模型被窃取的场景中会失效。窃取或泄露模型权重的攻击者可完全控制推理过程,只需运行未修改的采样器即可绕过水印,阻碍被盗后的归因。本研究中,我们提出了专家水印(Watermarking of Experts, WoE),一种利用稀疏混合专家(Mixture-of-Experts, MoE)模型独特结构特性的新型黑盒文本来源追踪方法。WoE对特定专家的词汇进行偏置,将水印信号嵌入转移至不可强制执行的推理包装器之外。该方法确保水印成为模型参数的固有属性,使防御者能够对被盗权重、泄露的检查点以及从被盗架构蒸馏出的次级稠密模型生成的文本进行归因,无需访问攻击者的部署或权重。我们在8个MoE模型上对WoE进行评估,结果显示其对可疑文本的水印检测成功,在1%的假阳性率下达到90.1%的平均真阳性率,最高可达94.9%,同时在很大程度上保留了模型的通用效用。此外,WoE在对抗性监督微调、模型提取和输出级 paraphrasing(意译)下仍可检测,迫使恶意行为者面临权衡:削弱归因信号需要额外的模型适配或文本重写操作,或损害结果输出的效用。
英文摘要
Large Language Model (LLM) watermarks provide a mechanism for text provenance, enabling model owners to identify machine-generated content and attribute it to a specific watermarked model. However, current LLM watermarking approaches predominantly rely on inference-time sampler methods and focus their analysis on dense models. Inference-time methods are only effective when the text is explicitly generated via the model owner's controlled API; they fail in a post-compromise scenario. An adversary who steals or leaks the model weights gains complete control over inference and can simply run an unmodified sampler, bypassing the watermark and preventing post-theft attribution. In this work, we introduce Watermarking of Experts (WoE), a novel black-box text provenance method that leverages the unique structural properties of sparse Mixture-of-Experts (MoE) models. WoE biases the vocabulary of specific experts and shifts the watermark signal embedding away from unenforceable inference wrappers. This approach ensures the watermark remains intrinsic to the model parameters, enabling defenders to attribute text generated by stolen weights, leaked checkpoints, and secondary dense models distilled from the stolen architecture without needing access to the adversary's deployment or weights. We evaluate WoE across eight MoE models, demonstrating successful watermark detection from suspect text, achieving an average true positive rate of 90.1% at a 1% false positive rate, reaching up to 94.9%, while largely preserving general model utility. Furthermore, WoE remains detectable under adversarial supervised fine-tuning, model extraction, and output-level paraphrasing, forcing malicious actors into a trade-off in which weakening the attribution signal requires additional model adaptation or text-rewriting operations, or compromises the utility of the resulting output.