发表机构
University of Chinese Academy of Sciences; Institute of Computing Technology, Chinese Academy of Sciences; Imperial College London(中国科学院大学; 中国科学院计算技术研究所; 伦敦帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SPIN通过局部随机轨迹邻域上的单步投影优化有界扰动,在不展开全轨迹的情况下破坏多种扩散编辑结果,在已见与未见指令下均全面超越现有方法。
AI 中文摘要
扩散模型极大地推进了指令引导的图像编辑,同时也引发了对未经授权图像操作的担忧。图像免疫通过向输入图像添加难以察觉的扰动来破坏后续编辑,从而应对这一风险。由于编辑请求在图像发布时是未知的,保护措施应在用于构建扰动的指令之外仍然有效。现有的免疫方法要么需要昂贵的全轨迹反向传播,要么使用中间目标,其效果可能被后续去噪削弱。同时,单一推理路径对替代去噪延续提供的反馈有限。为解决这些挑战,我们提出了SPIN,一种通过局部随机轨迹邻域上的单步投影进行图像免疫的框架。从早期去噪状态出发,SPIN在同一指令下生成随机相邻状态,并通过单步投影预测其干净潜变量,而无需完整展开。然后我们优化一个有界输入扰动,以最大化这些预测与干净编辑参考之间的平均偏差,从而促使扰动破坏多种可能的编辑结果。在两个图像编辑器上的实验表明,在已见指令和更具挑战性的未见指令设置下,SPIN在全部六项指标上均优于对比方法,保护性能显著提升。
英文摘要
Diffusion models have greatly advanced instruction-guided image editing, while also raising concerns about unauthorized image manipulation. Image immunization addresses this risk by adding imperceptible perturbations to an input image to disrupt subsequent edits. Since editing requests are unknown at image release, protection should remain effective beyond the instruction used to construct the perturbation. Existing immunization methods either require costly full-trajectory backpropagation or use intermediate objectives whose effects may be weakened by subsequent denoising. Meanwhile, a single inference path provides limited feedback about alternative denoising continuations. To address these challenges, we propose \textsc{SPIN}, a framework for image immunization via one-step projection over local stochastic trajectory neighborhoods. Starting from an early denoising state, \textsc{SPIN} generates stochastic neighboring states under the same instruction and predicts their clean latents through one-step projection without full unrolling. We then optimize a bounded input perturbation to maximize the average deviation of these predictions from a clean-edit reference, encouraging the perturbation to disrupt multiple possible editing outcomes. Experiments on two image editors demonstrate substantial gains in protection performance, with \textsc{SPIN} outperforming compared methods across all six metrics under seen instructions and in the more challenging unseen instruction setting.