arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12731cs.CV

GRACE:扩散模型中基于几何引导保留的自适应概念擦除

GRACE: Adaptive Concept Erasure with Geometry-Guided Retention in Diffusion Models

Qinghui Gong, Yihuai Liang, Yuanlun Xie, Deepak Kumar Jain, Vitomir Štruc, Zhengchun Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

GRACE提出结构化概念擦除框架,通过语义加权子空间估计、子空间约束适配器、自动解耦安全锚定和能量驱动动态门控,在扩散模型中实现局部干预,平衡擦除效果与生成保真度,显著提升NSFW降低率并降低FID。

中文摘要 AI 辅助

文本到图像(T2I)扩散模型不可避免地会从大规模预训练数据中内化敏感或不合规的概念,因此需要事后概念擦除。然而,现有的擦除方法通常缺乏对参数更新的显式约束,导致过度干预和意外的语义漂移。此外,许多方法依赖手工制作的反事实监督(例如替代提示),这带来了大量的数据构建成本,限制了其对新概念的可扩展性。为解决这些局限性,我们提出了GRACE,一个结构化概念擦除框架,旨在实现局部化和选择性干预。具体来说,我们引入了一种语义加权的敏感子空间估计来精确锁定干预方向,并采用轻量级子空间约束适配器以防止全局语义扰动。为了消除对人工提示工程的依赖,我们设计了一种自动解耦的安全锚定机制。为了缓解过度干预引起的语义漂移,我们引入了一种能量驱动的动态门控机制,在推理时自适应地控制干预的时机和强度。大量实验表明,我们的方法在擦除效果和生成保真度之间取得了优越的平衡。与五种最先进的(SOTA)概念擦除方法的平均性能相比,我们的方法将细粒度NSFW降低率提高了17.86%,同时将宏平均目标CLIP分数和面向保留的Fréchet Inception Distance(FID)分别降低了4.75%和50.58%,表明在显著改善原始模型生成效用的保留的同时,实现了更强的概念抑制。

英文摘要

Text-to-image (T2I) diffusion models inevitably internalize sensitive or non-compliant concepts from large-scale pretraining data, necessitating post-hoc concept erasure. However, existing erasure methods often lack explicit constraints on parameter updates, leading to over-intervention and unintended semantic drift. In addition, many methods rely on manually crafted counterfactual supervision, such as surrogate prompts, which incurs substantial data construction costs that limit scalability to new concepts. To address these limitations, we propose GRACE, a structured concept erasure framework designed to enable localized and selective intervention. Specifically, we introduce a semantically weighted sensitive subspace estimation to precisely lock intervention directions, and employ lightweight subspace-constrained adapters to prevent global semantic disturbance. To eliminate the dependency on manual prompt engineering, we design an automatically decoupled safe-anchor mechanism. To mitigate semantic drift induced by excessive intervention, we introduce an energy-driven dynamic gating mechanism that adaptively controls the timing and strength of intervention at inference. Extensive experiments demonstrate that our method achieves a superior balance between erasure effectiveness and generation fidelity. Compared with the average performance of five state-of-the-art (SOTA) concept erasure methods, our method improves the fine-grained NSFW reduction rate by $17.86\%$, while reducing the macro-averaged target CLIP Score and preservation-oriented Fréchet Inception Distance (FID) by $4.75\%$ and $50.58\%$, respectively, indicating stronger concept suppression with substantially improved preservation of the original model's generative utility.

发表机构

  • Southwest Jiaotong University(西南交通大学)
  • Chengdu University(成都大学)
  • Dalian University of Technology(大连理工大学)
  • University of Ljubljana(卢布尔雅那大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑