发表机构
ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何让开源文本语言模型水印在模型合并后仍具耐久性,提出合并对抗训练算法,该算法能将水印融入模型权重,在保留下游能力的同时优于所有基线,还评估了针对现实合并场景的水印,表明对抗训练可提高水印耐久性。
AI 中文摘要
开源语言模型(OSMs)性能接近最先进水平,此前通过嵌入水印算法追踪其生成文本。然而,OSMs会在训练后修改,模型合并会强烈移除水印。关键问题是如何使OSM水印在合并后仍存在。本文首次展示如何设计抗合并的OSM水印。提出合并对抗训练算法,将水印蒸馏到模型权重中且对后续合并稳健。该方法始终优于所有基线,还首次针对现实合并场景评估OSM水印,结果表明对抗训练能提高水印耐久性。
英文摘要
Open-source LLMs (OSMs)arereaching near state-of-the-art performance, prompting prior works to trace the text they generate by embedding text watermarking algorithms directly into their weights. Yet, OSMs are subject to post-training modifications, which has been shown to remove the watermark. Model merging in particular, a prominent method used for combining expert knowledge and preventing catastrophic forgetting, strongly removes such OSM watermarks. A key question is how to enable OSM watermarks that survive subsequent merging. In this work, we show for the first time how to design an OSM watermark that is durable against model merging. We propose Merge-Adversarial Training, an adversarial training algorithm to distill text watermarks into model weights while being robust to subsequent model merging. Our approach consistently outperforms all baselines (e.g. with SLERP up to +51 percentage points (pp) TPR@1%FPR with +25 pp on average) while preserving downstream capabilities. We also for the first time evaluate OSM watermarks against realistic merge scenarios, representing common use-cases such as combining expert capabilities or preventing catastrophic forgetting, and with 3 prominent merging algorithms. More broadly, our findings suggest that adversarial training is a reliable approach for increasing OSM watermark durability against post-training modifications.