arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10870cs.CV

NullEdit:通过视觉语言模型条件重定向实现隐秘图像保护

NullEdit: Stealthy Image Protection via VLM Condition Redirection

Weiyao Huang, Liqin Wang, Ziqi Sheng, Wei Lu

AI总结:

NullEdit针对VLM表示进行重定向,可隐秘抑制图像编辑且保留源内容,在基准模型上平均降低EditReward IF分数0.813,实现了高效的图像保护。

AI中文摘要:

现代图像编辑器将视觉语言模型(VLM)与扩散Transformer(DiT)骨干网络结合,可在无需微调的情况下根据指令修改单张参考图像,该能力也使得公开图像易被未经授权地篡改。现有推理时防御方法要么通过明显的损坏使编辑失效,从而暴露了保护措施;要么允许编辑进行,但存在身份或参考内容漂移,因此无法阻止编辑行为本身。我们转而针对一种隐秘且无害的“无操作”场景:抑制请求的编辑,输出保持自然且保留源图像,无明显伪影或身份替换,且恶意指令请求的有害语义不存在。我们提出NullEdit,其目标是在VLM表示(由参考图像和指令共同形成)作用于下游DiT骨干网络之前对其进行处理。利用正常编辑和无编辑锚点,NullEdit重定向该表示,同时通过跨提示梯度平均将保护扩展至未见过的指令。在CelebA-HQ和VGGFace2数据集上针对Step1X-Edit和Qwen-Image-Edit模型,NullEdit与SOTA基线相比平均将EditReward IF分数降低0.813,同时保留主体身份和源内容。

英文摘要:

Modern image editors combine vision-language models (VLMs) with diffusion transformer backbones to modify a single reference image according to instructions without fine-tuning. This capability also enables unauthorized manipulation of publicly released images. Existing inference-time defenses either invalidate edits through conspicuous corruption, thereby exposing the protection, or allow them to proceed with identity or reference content drift, thereby failing to prevent the editing behavior itself. We instead target a stealthy and harmless no-op in which the requested edit is suppressed, the output remains natural and source-preserving without conspicuous artifacts or identity replacement, and harmful semantics requested by malicious instructions are absent. We propose NullEdit, which targets the VLM representation jointly formed from the reference image and instruction before it conditions the downstream DiT backbone. Using normal-edit and no-edit anchors, NullEdit redirects this representation, while cross-prompt gradient averaging transfers protection to held out instructions. Across Step1X-Edit and Qwen-Image-Edit on CelebA-HQ and VGGFace2, NullEdit reduces the EditReward IF score by 0.813 on average relative to the SOTA baseline while preserving subject identity and source content.

补充信息

↑