arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VisualErase:用于文本到图像扩散模型中稳健概念擦除的双分支视觉轨迹重定向

VisualErase: Dual-Branch Visual Trajectory Redirection for Robust Concept Erasure in Text-to-Image Diffusion Models

Qianlong Xiang, Miao Zhang, Kun Wang, Yupeng Hu, Junhui Hou, Liqiang Nie

arXiv 2610.05000首次发表:更新:

发表机构

Harbin Institute of Technology (Shenzhen); City University of Hong Kong; Shenzhen Loop Area Institute; National University of Singapore; Shandong University(哈尔滨工业大学(深圳); 香港城市大学; 深圳河套学院; 新加坡国立大学; 山东大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VisualErase通过双分支视觉轨迹重定向,将概念相关的生成轨迹导向概念移除结果,在风格、名人、裸体擦除中显著降低攻击成功率,同时保持生成质量。

AI 中文摘要

概念擦除对于文本到图像扩散模型的安全部署至关重要,因为这些模型可能会重现从无约束的大规模数据中学习到的有害、受版权保护或隐私敏感的内容。现有方法通常通过重定向与目标相关的文本到图像映射来擦除不需要的概念,同时保留通用生成能力。然而,近期研究表明,被擦除的模型可能仍然保留目标概念的视觉生成轨迹,使其容易受到对抗性恢复攻击,并揭示了重定向文本到图像映射与真正移除视觉知识之间的根本差距。为弥合这一差距,我们提出了VisualErase,一种新范式,将携带概念的视觉生成轨迹重定向到明确定义的概念移除结果。为实现这种重定向,我们使用保持结构的图像编辑来为源图像构建内容对齐、概念移除的对应图像,提供保留非目标内容的明确视觉端点。然后,我们从每个源图像-对应图像对中推导出一个去噪目标,并使用双分支重定向损失来使文本条件和无条件预测均与该目标对齐,因为仅靠条件监督无法显式约束无文本引导的生成。为减轻概念擦除对非目标生成的负面影响,我们联合优化重定向损失与对应保留损失,后者匹配冻结预训练模型的去噪预测。在风格、名人和裸体擦除任务中,VisualErase将七种攻击的最大攻击成功率分别限制为0%、8%和0.1%,同时保持通用生成质量。这些结果凸显了视觉轨迹重定向对于超越仅文本到图像映射的稳健概念擦除的重要性。

英文摘要

Concept erasure is essential for the safe deployment of text-to-image diffusion models, as they may reproduce harmful, copyrighted, or privacy-sensitive content learned from unconstrained large-scale data. Existing methods typically erase unwanted concepts while preserving general generation capability by redirecting target-related text-to-image mappings. However, recent studies show that erased models may still retain visual generative trajectories of target concepts, leaving them vulnerable to adversarial recovery attacks and revealing a fundamental gap between redirecting text-to-image mappings and truly removing visual knowledge. To bridge this gap, we propose VisualErase, a new paradigm that redirects concept-bearing visual generative trajectories toward explicitly defined concept-removed outcomes. To enable this redirection, we use structure-preserving image editing to construct content-aligned, concept-removed counterparts for source images, providing explicit visual endpoints that retain non-target content. We then derive a denoising target from each source-to-counterpart pair and use a dual-branch redirection loss to align both text-conditioned and unconditional predictions with this target, since conditional supervision alone does not explicitly constrain generation without textual guidance. To mitigate the adverse effects of concept erasure on non-target generation, we jointly optimize the redirection loss with a counterpart retention loss that matches denoising predictions from the frozen pretrained model. Across style, celebrity, and nudity erasure, VisualErase limits the maximum attack success rate over seven attacks to 0%, 8%, and 0.1%, respectively, while retaining general generation quality. These results highlight the importance of visual trajectory redirection for robust concept erasure beyond text-to-image mappings alone.

CommentsThe project page is https://qianlong0502.github.io/VisualErase-Homepage

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑