arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

剔除坏种子:文本到图像扩散模型的初始噪声鲁棒性遗忘

Weeding Out Bad Seeds: Initial-Noise-Robust Unlearning for Text-to-Image Diffusion Models

Arian Komaei Koma, Seyed Amir Kasaei, Aida Aryafar, Matin Ghiasi, Ali Aghayari, Amirhossein Souri, Mohammad Mosayyebi, AmirMahdi Sadeghzadeh, Mohammad Hossein Rohban

arXiv 2609.37537首次发表:更新:

发表机构

Sharif University of Technology; Hong Kong University of Science and Technology(谢里夫理工大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对文本到图像扩散模型遗忘对初始噪声敏感的问题,提出自适应概念条件采样策略,动态聚焦梯度更新,显著降低概念重现率并增强对抗鲁棒性。

AI 中文摘要

机器遗忘已成为一种关键的后期安全措施,用于在不进行昂贵重训练的情况下从文本到图像(T2I)模型中擦除敏感概念。然而,我们发现当前最先进(SOTA)的方法因严重缺乏对噪声初始化的鲁棒性而显得脆弱。我们将此现象称为“概率性遗忘”:被抑制的概念在特定的随机初始噪声条件下会重新出现,尽管在其他初始化下看似已被遗忘。我们将此失败归因于遗忘过程中标准高斯采样与遗忘目标之间的错位。由于目标概念仅在遗忘阶段的特定初始噪声区域中显现,均匀随机采样产生稀疏、信息量不足的梯度更新,无法实现稳健的擦除。为解决此问题,我们提出一种自适应的、基于概念条件的采样策略,该策略动态地将梯度更新集中在目标概念显现的区域,并对信息量不足的区域进行降权。我们将我们的框架与六种不同的SOTA遗忘方法集成,涵盖四种扩散骨干网络,并在安全性、物体和艺术风格遗忘以及黑盒和白盒对抗攻击下进行评估。我们的方法在四个基线上,将随机初始化下的条件性裸体重现率平均降低了67.2%,并在两种对抗评估中降低了攻击成功率。跨概念领域,自适应噪声采样增强了对抗鲁棒性和非目标保留,同时保持了有竞争力的生成质量和目标擦除性能。

英文摘要

Machine unlearning has emerged as a critical post-hoc safety measure to erase sensitive concepts from Text-to-Image (T2I) models without prohibitive retraining. However, we reveal that current state-of-the-art (SOTA) approaches are brittle due to a severe lack of robustness to noise initialization. We call this phenomenon ``probabilistic forgetting'': suppressed concepts re-emerge under specific random initial noise conditions, despite appearing unlearned on other initializations. We trace this failure to the misalignment between standard Gaussian sampling during unlearning and the unlearning objective. Since the target concept manifests only in specific initial noise regions throughout the unlearning phase, uniform random sampling yields sparse, uninformative gradient updates that fail to drive robust erasure. To overcome this issue, we propose an adaptive, concept-conditioned sampling strategy that dynamically concentrates gradient updates on regions where the target concept manifests, down-weighting uninformative areas. We integrate our framework with six distinct SOTA unlearning methods across four diffusion backbones and evaluate it across safety, object, and artistic-style unlearning, as well as under black-box and white-box adversarial attacks. Our method reduces the conditional nudity re-emergence rate across random initializations by 67.2% on average over four baselines and lowers attack success rates across both adversarial evaluations. Across concept domains, Adaptive Noise Sampling strengthens adversarial robustness and non-target retention while preserving competitive generative quality and target-erasure performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑