RAPID:针对文本到图像服务未授权模型蒸馏的实时防御
RAPID: A Real-Time Defense Against Unauthorized Model Distillation for Text-to-Image Services
浏览论文内容
中文总结 AI 辅助
针对T2I服务易受黑盒蒸馏攻击的问题,提出RAPID自引用潜在最大化框架,集成防御扰动于VAE解码器,实现实时保护,在四个模型和数据集上显著降低替代模型质量且保持视觉保真度。
中文摘要 AI 辅助
基于扩散的文本到图像(T2I)模型越来越多地被用于视觉内容创作,使其生成能力成为宝贵的知识产权资产。然而,这种能力容易受到基于黑盒输出的蒸馏攻击,攻击者查询服务、收集提示-图像对,并训练一个模仿其生成行为的未授权替代模型。现有的基于扰动的防御方法采用样本级优化,使生成的图像对未授权训练具有破坏性,但引入了大量计算和延迟,降低了在线T2I服务的可用性。一个自然的解决方案是将防御性扰动集成到VAE解码器中,使受保护模型直接生成防御图像,而无需在线样本级优化。然而,现有的样本级目标难以迁移到共享解码器设置中。我们通过实验发现,防御性共享解码器引起的潜在空间偏移远小于样本级优化,这表明在该设置中,目标可达性比破坏性更重要。为克服这一限制,我们提出RAPID,一种自引用潜在最大化框架,它消除外部依赖,并鼓励相同的模型更新在训练样本间产生一致的破坏性效果,从而提高可达性。我们进一步引入重建引导的颜色正则化,以阻断潜在捷径并增强视觉破坏。在四个T2I模型和四个数据集上的大量实验,与五个代表性基线的比较表明,RAPID持续降低替代模型的生成质量,同时保持服务的视觉保真度。我们的工作为部署的T2I系统中针对未授权蒸馏的实时保护建立了一种范式。
英文摘要
Diffusion-based text-to-image (T2I) models are increasingly used for visual content creation, making their generation capability a valuable intellectual property asset. However, this capability is vulnerable to black-box output-based distillation, where an adversary queries the service, collects prompt-image pairs, and trains an unauthorized substitute model that mimics its generation behavior. Existing perturbation-based defenses apply sample-wise optimization to make generated images disruptive to unauthorized training, but introduce substantial computation and latency that reduce the usability of online T2I services. A natural solution is to integrate defensive perturbations into the VAE decoder, allowing the protected model to generate defended images directly without online sample-wise optimization. However, existing sample-wise objectives struggle to transfer to the shared decoder setting. We empirically find that a defensive shared decoder induces a substantially smaller latent shift than sample-wise optimization, suggesting that objective reachability matters more than destructiveness in this setting. To overcome this limitation, we propose RAPID, a self-referenced latent maximization framework that removes external dependencies and encourages the same model update to induce consistently disruptive effects across training samples, thereby improving reachability. We further introduce reconstruction-guided color regularization that blocks the latent shortcut and reinforces visual disruption. Extensive experiments on four T2I models and four datasets, with comparisons against five representative baselines, show that RAPID consistently degrades substitute-model generation quality while preserving service visual fidelity. Our work establishes a paradigm for real-time protection against unauthorized distillation in deployed T2I systems.
发表机构
- University of Electronic Science and Technology of China(电子科技大学)
- Nanyang Technological University(南洋理工大学)
- City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。