认证文本到图像扩散模型中的概念遗忘
Certifying Concept Unlearning in Text-to-Image Diffusion Models
- Imperial College London(伦敦帝国理工学院)
- TU Wien(维也纳工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对T2I扩散模型概念遗忘评估依赖攻击成功率而忽视残余泄漏的问题,提出结合统计认证与最坏情况分析的认证框架,在三大概念类别和六种方法上验证,泄漏界超过攻击成功率16.2%,证明认证是可靠审计的必要补充。
AI中文摘要:
现有的文本到图像(T2I)扩散模型中概念遗忘的评估主要依赖于通过自动化对抗性提示搜索获得的攻击成功率。然而,这些指标仅提供了在有限查询集上的经验证据,而将更广泛提示空间中的残余泄漏在很大程度上未量化。这一局限性可能导致高估遗忘效果并低估安全风险。为解决这一差距,我们引入了一种新颖的T2I概念遗忘认证框架,该框架在残余概念泄漏上提供具有有界误差的高置信度保证。我们的方法将统计认证与沿概念相关嵌入方向的 worst-case 分析相结合,以在用户指定的置信水平下推导泄漏概率的显式上界。我们在三大概念类别(即 NSFW 内容、艺术风格和名人身份)以及六种最先进的遗忘方法上评估了我们的框架。认证的泄漏界始终超过标准攻击成功率 16.2%,揭示了现有评估协议遗漏的大量残余风险。至关重要的是,我们的结果表明,基于经验攻击的评估可能显著低估残余泄漏,并确立认证作为可靠审计 T2I 扩散模型中概念遗忘的必要补充。
英文摘要:
Existing evaluations of concept unlearning in text-to-image (T2I) diffusion models primarily rely on attack success rates obtained through automated adversarial prompt search. However, these metrics provide only empirical evidence over a finite set of queries and leave residual leakage over the broader prompt space largely unquantified. This limitation can lead to overestimating unlearning effectiveness and underestimating safety risks. To address this gap, we introduce a novel certification framework for T2I concept unlearning that provides high-confidence guarantees with bounded error on residual concept leakage. Our approach combines statistical certification with worst-case analysis along concept-relevant embedding directions to derive explicit upper bounds on leakage probability under user-specified confidence levels. We evaluate our framework across three major concept categories namely NSFW content, artistic styles, and celebrity identities, and six state-of-the-art unlearning methods. Certified leakage bounds consistently exceed standard attack success rates by 16.2%, uncovering substantial residual risks missed by existing evaluation protocols. Crucially, our results demonstrate that empirical attack-based evaluations can significantly underestimate residual leakage and establish certification as a necessary complement for reliable auditing of concept unlearning in T2I diffusion models.