文本到图像模型的持久水印
Persistent Watermarking of Text-to-Image Models
查看机构详情
- University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对文本到图像模型,提出对比式水印目标,嵌入持久触发数据,在多种下游修改和攻击下保持高检测率,接近100% TPR@FPR<10^{-4}。
中文摘要 AI 辅助
文本到图像(T2I)生成在公众中日益流行,鉴于其昂贵的训练成本,推动了可靠版权保护机制的发展。对手可能未经授权获取并重用预训练的T2I模型,然后通过API服务提供修改版本。这些修改可能源于普通的下游适应或故意试图消除所有权,包括输入提示预处理、模型微调和输出后处理。从模型所有者的角度来看,一个关键挑战因此是嵌入触发数据,使其在此类修改下保持持久,同时保留模型正常的图像生成能力。在这项工作中,我们提出了一种对比式水印目标,其中包含一个明确鼓励水印模型在触发输入上表现不同于原始模型的项。实验表明,在广泛的下游修改和故意削弱水印的尝试中,触发数据的持久性显著强于先前方法,导致更高的检测率,通常接近100%的TPR@FPR<10^{-4}。
英文摘要
Text-to-image (T2I) generation is gaining increasing popularity with the general public, motivating the development of reliable mechanisms for copyrighting such models given their expensive training costs. An adversary may obtain and reuse a pretrained T2I model without authorization, and then serve a modified version through an API service. Such modifications may arise from ordinary downstream adaptation or deliberate attempts to erase ownership, including input-prompt preprocessing, model fine-tuning, and output post-processing. From the model owner's perspective, a key challenge is therefore to embed trigger data that remain persistent under such changes while preserving the model's normal image-generation capabilities. In this work, we propose a contrastive-style watermarking objective with a term that explicitly encourages the watermarked model to behave differently from the original model on trigger inputs. Experiments show substantially stronger trigger-data persistence than prior methods across a wide range of downstream modifications and deliberate attempts to weaken the watermark, resulting in higher detection rates, often approaching 100% TPR@FPR<$10^{-4}$.