arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05401cs.LGcs.AI

无概念能逃脱审计:面向可验证概念擦除的审计感知机器遗忘方法

No Concept Escapes the Audit: Auditing-Aware Unlearning for Verifiable Concept Erasure in Diffusion Models

Kaiyuan Deng, Yuchen Li, Gen Li, Yang Xiao, Geng Yuan, Xiaoyong Yuan, Bo Hui, Xiaolong Ma

首次发表
浏览论文内容

中文总结 AI 辅助

针对扩散模型概念擦除后仍可从内部表示恢复的问题,提出审计感知的AVCE框架,通过潜空间审计与锚点编辑实现可验证擦除,显著降低攻击成功率并提升审计分数。

中文摘要 AI 辅助

文本到图像扩散模型能够生成被禁止的内容,这促使通过机器遗忘技术实现概念擦除。大多数擦除方法在文本接口层面进行干预,通过提示修改或对文本条件权重的局部更新来实现,并通过模型对给定提示的输出进行评估。这种评估无法看到网络内部仍然编码的内容。潜空间审计绕过文本条件,直接探测去噪网络,表明被擦除的概念仍可从内部表示中恢复。我们发现,这一现象同样适用于那些旨在抵御对抗性提示的方法,并且问题随着被擦除概念数量的增加而加剧。我们提出了审计感知的扩散模型可验证概念擦除框架(AVCE),该框架将擦除建立在模型的潜表示之上。AVCE审计每个概念的嵌入邻域,并将发现的可攻击方向压缩为最弱几何点处的锚点。它在此锚点处以闭式形式编辑交叉注意力和自注意力投影,然后使用通路级审计损失对两条通路进行微调,利用正交梯度投影整合多个概念。在SD v1.5、SDXL和Flux 1.0上针对物体、露骨内容和艺术风格遗忘的实验表明,与最强基线相比,AVCE将攻击成功率降低了5.07倍,审计分数提高了3.84倍,同时保持了有竞争力的生成质量。

英文摘要

Text-to-image diffusion models can generate prohibited content, which motivates concept erasure through machine unlearning. Most erasure methods intervene at the text interface, through prompt modification or localized updates to text-conditioning weights, and they are evaluated by what the model outputs for given prompts. Such evaluation cannot see what the network still encodes. Latent-space auditing, which bypasses text conditioning and probes the denoising network directly, shows that erased concepts remain recoverable from internal representations. We find that this also holds for methods built to be robust against adversarial prompts, and that the problem grows with the number of erased concepts. We propose Auditing-Aware Unlearning for Verifiable Concept Erasure in Diffusion Models (AVCE), a framework that grounds erasure in the model's latent representations. AVCE audits the embedding neighborhood of each concept and condenses the discovered vulnerable directions into an anchor at the weakest geometric point. It edits cross-attention and self-attention projections in closed form at this anchor, then fine-tunes the two pathways with pathway-level auditing losses, using orthogonal gradient projection to consolidate multiple concepts. Experiments on SD v1.5, SDXL, and Flux 1.0 across object, explicit-content, and artistic-style unlearning show that AVCE reduces attack success rates by 5.07x and improves auditing scores by 3.84x over the strongest baseline, while preserving competitive generation quality.

发表机构

  • Clemson University(克莱姆森大学)
  • The University of Tulsa(塔尔萨大学)
  • University of Georgia(佐治亚大学)
  • The University of Arizona(亚利桑那大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑