arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

先学习,后记忆:扩散模型的谱偏置

First Learn, Then Memorize: The Spectral Bias of Diffusion Models

Raphaël Urfin, Tony Bonnaire, Giulio Biroli, Marc Mézard

arXiv 2609.35377首次发表:更新:

发表机构

École normale supérieure; Université PSL; Sorbonne Université; Université Paris Cité; Université Paris-Saclay; Bocconi University(巴黎高等师范学院; 巴黎文理研究大学; 索邦大学; 巴黎西岱大学; 巴黎萨克雷大学; 博科尼大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究揭示了扩散模型先泛化后记忆的机制,归因于NTK格拉姆矩阵谱的双主体结构,并通过解析与实验验证,提出截断秩或L2惩罚可调控泛化-记忆平衡。

AI 中文摘要

在有限数据集上训练的扩散模型首先学会生成新颖、高质量的样本,而仅在很久之后才坍缩到其训练集上。我们识别出这种时间尺度分离背后的机制以及探测该机制的客体。得分函数的训练动力学——在任意宽度下精确地——由在含噪训练数据上评估的神经正切核(NTK)的格拉姆矩阵所支配,因此泛化和记忆的时间尺度必定编码在其谱中。我们证明了这一点,并指出所涉及的结构在标准核设置中没有对应物。在得分匹配损失中,每个样本使用多个噪声实现(在固定噪声水平下的$m$个加噪副本)正是将格拉姆矩阵谱重构为两个不同部分的原因。第一部分,具有大特征值,携带目标分布的全局特征,且在$m=1$时已存在。第二部分,由重复加噪产生,由最小特征值组成,并支撑在与样本特定噪声方向对齐的特征向量上;它设定了一个在训练集规模$n$上参数化更大的记忆时间尺度。我们在两个方面建立了这一图景。在解析上,我们在懒惰高维极限下求解了线性($n \asymp d$)和多项式($n \asymp d^k$)样本复杂度下的谱,并通过偏差-方差分解证明,第一部分主体最小化逼近误差,而第二部分驱动与记忆相关的误差。在经验上,我们在CelebA上的卷积NTK以及在远超懒惰机制训练的有限宽度U-Net中展示了相同的双主体结构,并建立了因果联系:将格拉姆矩阵截断至秩$r$可调节泛化-记忆转变,而针对第二主体的$L_2$惩罚可抑制特征学习U-Net中的记忆。

英文摘要

Diffusion models trained on a finite dataset first learn to generate novel, high-quality samples and only much later collapse onto their training set. We identify the mechanism behind this separation of timescales and the object that probes it. The training dynamics of the score function are governed---exactly, and at any width---by the Gram matrix of the Neural Tangent Kernel (NTK) evaluated on the noisy training data, so the timescales of generalization and of memorization must be encoded in its spectrum. We show that they are, and that the structure responsible has no analogue in standard kernel settings. The use of multiple noise realizations per sample ($m$ noised copies at a fixed noise level) in the score-matching loss is what restructures the Gram matrix spectrum into two distinct parts. The first, of large eigenvalues, carries the global features of the target distribution and is present already for $m=1$. The second, which the repeated noising creates, consists of the smallest eigenvalues and is supported on eigenvectors aligned with the sample-specific noise directions; it sets a memorization timescale parametrically larger in the training set size $n$. We establish this picture on two fronts. Analytically, we solve the spectrum in the lazy high-dimensional limit for both linear ($n \asymp d$) and polynomial ($n \asymp d^k$) sample complexities, and prove through a bias--variance decomposition that the first bulk minimizes the approximation error while the second drives the error associated with memorization. Empirically, we show the same two-bulk structure in Convolutional NTKs on CelebA and in finite-width U-Nets trained well beyond the lazy regime, and we make the link causal: truncating the Gram matrix at rank $r$ tunes the generalization--memorization transition, and an $L_2$ penalty targeting the second bulk suppresses memorization in feature-learning U-Nets.

Comments53 pages, 13 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑