arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你的遗忘会暴露你:识别扩散模型中被擦除的概念

Your Unlearning Gives You Away: Identifying Erased Concepts in Diffusion Models

Kaiyuan Deng, Yuchen Li, Yang Xiao, Bo Hui, Geng Yuan, Xiaolong Ma

arXiv 2610.05601首次发表:更新:

发表机构

The University of Arizona; Peking University; The University of Tulsa; University of Georgia(亚利桑那大学; 北京大学; 塔尔萨大学; 佐治亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Tracer框架,通过权重足迹的频谱分析和足迹覆盖目标,无需生成图像即可快速识别扩散模型中被擦除的概念并估计其数量,较现有方法提速数百至数十万倍且准确率更高。

AI 中文摘要

现有的针对遗忘扩散模型的攻击假设被擦除的概念是预先已知的,并专注于恢复这些概念。然而,在实践中,模型提供者可能不会披露哪些概念已被移除,即使能够访问原始基础模型,攻击者也可能缺乏明确的攻击目标。在本文中,我们旨在回答以下关键但被忽视的问题:哪些概念已从模型中被擦除,以及总共擦除了多少个?为此,我们提出了Tracer,一个能够快速且准确地识别被擦除概念并估计其数量的框架。Tracer无需生成和分类图像即可高效识别被擦除的概念。通过结合权重足迹的轻量级频谱分析,它能够在大型候选词汇表上进行高效搜索。为了区分多个被擦除的概念,我们引入了一个足迹覆盖目标,用于指导顺序发现。Tracer通过检测候选置信度在所选概念覆盖擦除足迹时的急剧下降来估计被擦除概念的数量,无需标记示例进行校准。该框架仅需要轻量级线性代数和有限的向前探测,且无需事先了解遗忘算法。在文本到图像和文本到视频骨干网络以及多种遗忘方法上的实验表明,Tracer能够在几秒钟内识别被擦除的概念并估计其数量,在图像和视频模型上分别比MIA和暴力搜索实现了150至137,000倍和133至20,000倍的加速,同时具有显著更高的识别准确率。

英文摘要

Existing attacks on unlearned diffusion models assume that the erased concepts are known in advance and focus on recovering them. In practice, however, model providers may not disclose which concepts have been removed, and even with access to the original base model, an adversary may still lack a clear target to attack. In this paper, we aim to answer the following critical but overlooked questions: which concepts have been erased from the model, and how many have been erased in total? To this end, we present Tracer, a framework that rapidly and accurately identifies erased concepts and estimates their number. Tracer efficiently identifies erased concepts without generating and classifying images. By combining lightweight spectral analysis of weight footprints, it enables efficient search over large candidate vocabularies. To distinguish multiple erased concepts, we introduce a footprint coverage objective that guides sequential discovery. Tracer estimates the number of erased concepts by detecting a sharp decline in candidate confidence as the selected concepts account for the erasure footprint, without requiring labeled examples for calibration. The framework requires only lightweight linear algebra and limited forward probes, with no prior knowledge of the unlearning algorithm. Experiments across text-to-image and text-to-video backbones and diverse unlearning methods demonstrate that Tracer identifies erased concepts and estimates their number in seconds, achieving 150 to 137,000 times and 133 to 20,000 times speedups over MIA and brute-force search on image and video models, respectively, with substantially higher identification accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑