发表机构
Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi; Indian Institute of Technology Delhi; Jawaharlal Nehru Centre for Advanced Scientific Research; TCS Research Labs(德里印度理工学院亚迪人工智能学院; 德里印度理工学院; 贾瓦哈拉尔·尼赫鲁先进科学研究中心; 塔塔咨询服务公司研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现扩散模型的记忆在生成样本前就编码于能量景观几何,通过循环去噪可探测潜在记忆,在Stable Diffusion v1.4上实现了高AUC和TPR的记忆提示检测。
AI 中文摘要
扩散模型在训练早期表现出泛化能力,后期会复现单个训练样本。标准测试仅在单次生成产生近乎副本时才检测到记忆,导致发布的模型在输出失效前无法被审计。我们证明记忆在生成样本出现前就已编码在学习到的能量景观几何中,这种状态称为潜在记忆。利用分数散度和盆地体积,我们发现局部盆地在第一个记忆样本出现前就围绕训练样本形成,并将其与保留样本分离,其起始遵循与记忆时间相同的O(n)缩放规律。我们通过循环去噪探测这些盆地,该方法会反复应用部分加噪和去噪。在精确经验分数下,我们证明从孤立训练样本附近开始的循环会以高概率在任意有限次循环中恢复该样本并返回。在训练模型中,循环可从CelebA和CIFAR-10的检查点中恢复训练图像,这些检查点的单次生成样本不含副本;在CelebA的一个检查点(单次副本率为0.1%)上,500次循环将记忆分数提升至30%以上。循环还会揭示与单个训练图像不匹配的退化吸引子,且随训练进行而消失,因此驻留在盆地本身并不意味着记忆。这些发现适用于高斯混合、CelebA和CIFAR-10,涵盖优化器、架构、噪声调度和训练集大小,并扩展到现成的Stable Diffusion v1.4,其中循环的条件-无条件散度 gap 以0.944的AUC和1% FPR下0.866的TPR将记忆提示与非记忆提示分离。更广泛地说,扩散模型记忆了什么是其学习分布的几何和稳定性的属性,评估它需要检查该结构而非仅生成输出。
英文摘要
Diffusion models generalize early in training and later reproduce individual training samples. Standard tests detect memorization only once one-shot generation produces near-copies, leaving a released model unaudited until its outputs fail. We show that memorization is encoded in the geometry of the learned energy landscape before it appears in generated samples, a state we call latent memorization. Using score divergence and basin volume, we find that localized basins form around training samples and separate them from held-out samples before the first memorized sample appears, with an onset that follows the same $O(n)$ scaling as the memorization time. We probe these basins with cyclic denoising, which repeatedly applies partial noising and denoising. Under the exact empirical score, we prove that cycling started near an isolated training sample recovers it and returns to it over any finite number of cycles with high probability. In trained models, cycling recovers training images from CelebA and CIFAR-10 checkpoints whose one-shot samples contain no copies, and at a CelebA checkpoint with 0.1% one-shot copies, 500 cycles raise the memorized fraction above 30%. Cycling also reveals degenerate attractors that match no single training image and fade as training proceeds, so residence in a basin does not by itself imply memorization. These findings hold on a Gaussian mixture, CelebA, and CIFAR-10 across optimizers, architectures, noise schedules, and training-set sizes, and extend to off-the-shelf Stable Diffusion v1.4, where the cycled conditional-unconditional divergence gap separates memorized from non-memorized prompts with an AUC of 0.944 and a TPR of 0.866 at 1% FPR. More broadly, what a diffusion model has memorized is a property of the geometry and stability of its learned distribution, and assessing it requires examining this structure rather than generated outputs alone.
Comments42 pages, 24 figures, 7 tables