arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

概念分数再学习:针对概念擦除的统一跨架构攻击

Concept Score Relearning: A Unified Cross-Architecture Attack on Concept Erasure

Hong Xi Tae, Jiaming Zhang, Wenwen He, Xuan Wang, Wei Yang Bryan Lim

arXiv 2609.33445首次发表:更新:

发表机构

College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出统一参数级框架CSR,在U-Net和Transformer两种架构中通过原生预测空间优化再激活被擦除概念,无需外部数据集,跨架构实验验证其有效性,严格裸体场景下ASR分别达50.47%和40.29%。

AI 中文摘要

概念擦除旨在抑制文本到图像生成模型中的不良知识。然而,现有的鲁棒性评估通常依赖于针对特定模型架构定制的再学习攻击。我们研究了两种截然不同的生成范式中的概念再激活:噪声预测U-Net和流匹配Transformer。我们引入了概念分数再学习(CSR),这是一个统一的参数级框架,通过在各自的原生预测空间内优化每个模型来再激活被擦除的概念。CSR不需要外部目标概念图像数据集,并将相同的概念导向目标应用于基于U-Net的Stable Diffusion和基于Transformer的FLUX。跨多种概念和多种擦除方法的实验表明,两种架构均能实现一致的概念再激活,凸显了CSR的跨架构适用性以及看似被擦除的概念的持续可恢复性。对于严格的裸体内容,CSR在FLUX上达到50.47%的平均攻击成功率(ASR),在Stable Diffusion上达到40.29%,在所有评估的安全设置中均排名第一。

英文摘要

Concept erasure aims to suppress undesirable knowledge in text-to-image generative models. However, existing robustness evaluations typically rely on relearning attacks tailored to specific model architectures. We study concept reactivation across two substantially different generative paradigms: noise-prediction U-Nets and flow-matching Transformers. We introduce \textbf{Concept Score Relearning (CSR)}, a unified parameter-level framework that reactivates erased concepts by optimizing each model within its native prediction space. CSR requires no external target-concept image dataset and applies the same concept-directed objective to both U-Net-based Stable Diffusion and Transformer-based FLUX. Experiments across diverse concepts and multiple erasure methods demonstrate consistent concept reactivation across both architectures, highlighting the cross-architecture applicability of CSR and the persistent recoverability of apparently erased concepts. For strict nudity, CSR reaches average ASRs of 50.47\% on FLUX and 40.29\% on Stable Diffusion, consistently ranking first across all evaluated safety settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑