arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

擦除还是不擦除:基于保留感知自适应排序子空间扩展的无训练鲁棒概念擦除

To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion

Shaswati Saha, Rajasekhar Anguluri, Manas Gaur

arXiv 2607.23492首次发表:更新:

发表机构

University of Maryland Baltimore County(马里兰大学巴尔的摩县分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对文本到图像扩散模型的概念擦除技术面临的鲁棒性与效用权衡问题,提出无训练框架PARSE,通过自适应发现擦除与保留概念、编辑交叉注意力值空间及迭代搜索触发因素来实现鲁棒概念擦除,引入BEUS评估,实验证明其有效性。

AI 中文摘要

概念擦除技术(CETs)编辑文本到图像的扩散模型以擦除不期望的目标,如NSFW内容或受版权保护的风格,同时保留模型在良性概念上的效用。当前的CETs在擦除鲁棒性和效用之间面临权衡。许多CETs依赖手动指定、由语言模型生成或通过CLIP图像-文本相似性选择的静态概念库,无法对去噪过程中提示如何引导模型进行建模。我们提出了保留感知自适应排序子空间扩展(PARSE),这是一个用于潜在扩散模型中鲁棒概念擦除的无训练框架。给定目标,PARSE通过无分类器指导查询扩散模型,动态发现目标诱导的擦除概念和附近的保留概念。然后,它通过保留感知投影编辑交叉注意力值空间,去除目标方向同时保留保留方向。对于超出此词汇索引空间的触发因素,PARSE通过文本反转迭代搜索重新出现的触发因素,并仅在新的触发方向与保留语义不冲突时自适应扩展擦除子空间。我们还引入了平衡擦除效用分数(BEUS),通过有界单调变换和谐波均值聚合结合鲁棒性(多次攻击下的ASR)和效用保留(FID)。实验表明,PARSE能稳健地擦除多个概念而不牺牲编辑后的效用。

英文摘要

Concept erasure techniques (CETs) edit text-to-image diffusion models to erase undesired targets such as NSFW content or copyrighted styles, while preserving model utility on benign concepts. Current CETs face a trade-off between erasure robustness and utility: stronger edits erase the target more reliably but degrade utility on non-target concepts, and vice versa. This stems from how existing methods define what to erase and what to preserve. Many CETs rely on static concept banks specified manually, generated by LLMs, or selected by CLIP image-text similarity. Such banks do not model how prompts steer the model during denoising, leaving it vulnerable to triggers that reintroduce the target while suppressing nearby benign concepts. We present Preservation-aware Adaptive Ranked Subspace Expansion (PARSE), a training-free framework for robust concept erasure in latent diffusion models. Given a target, PARSE queries the diffusion model with classifier-free guidance to dynamically discover target-inducing erase concepts and nearby retain concepts in the model vocabulary. It then edits the cross-attention value space with a preservation-aware projection that removes target directions while leaving retain directions intact. For triggers beyond this vocabulary-indexed space, PARSE iteratively searches for re-emergence triggers by textual inversion and adaptively expands the erased subspace only when a new trigger direction does not conflict with retain semantics. We also introduce the Balanced Erasure Utility Score (BEUS), which combines robustness (ASR under multiple attacks) and utility preservation (FID) via bounded monotone transforms and harmonic mean aggregation. Experiments on NSFW, artistic style, and object erasure, with a large-scale robustness-utility analysis over many CET baselines, show that PARSE erases multiple concepts robustly without sacrificing post-edit utility.

CommentsAccepted to ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑