arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩散模型中的双重下降与恶性过拟合

Double Descent and Malign Overfitting in Diffusion Models

Raphaël Urfin, Tony Bonnaire, Giulio Biroli, Marc Mézard

arXiv 2609.26392首次发表:更新:

发表机构

Laboratoire de Physique de l’École normale supérieure, ENS, Université PSL, CNRS, Sorbonne Université, Université Paris Cité; Université Paris-Saclay, CNRS, Institut d’Astrophysique Spatiale; Bocconi University(巴黎高等师范学院物理实验室,ENS,PSL大学,CNRS,索邦大学,巴黎西岱大学; 巴黎-萨克雷大学,CNRS,空间天体物理研究所; 博科尼大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过U-Net实验与随机特征理论,揭示扩散模型训练中过参数化导致恶性过拟合(记忆化)而非良性过拟合的机制,并证明结合正则化(岭惩罚或早停)的过参数化仍能带来性能提升。

AI 中文摘要

深度学习中的传统观点认为,过参数化——即参数数量$p$多于训练样本数量$n$——是有益的:更大的模型泛化能力更好,即使没有正则化,插值模型也能很好地泛化,测试误差遵循双重下降曲线。人们可能期望扩散模型也具有同样的良性过拟合,因为其训练可归结为回归问题,即最小化二次得分匹配损失。然而,观察到的现象恰恰相反:这里的过拟合是灾难性的,将模型推向记忆化状态。我们通过结合在CelebA上训练的U-Net实验与随机特征模型(我们为其推导出闭式学习曲线)来解决这一悖论。我们表明,当每个训练样本的噪声实现数量固定为$m$时,插值峰确实会出现,但出现在$p\sim nm$处,而非标准回归中的$p\sim n$处。然而,测试损失的上升要早得多,在$p\sim n$处就开始,与$m$无关。这种过拟合是恶性的,因为尽管训练的隐式正则化完全发挥作用,但它将模型推向经验得分,该得分记忆了训练集,而非真实得分。偏差-方差分解揭示了其机制:得分估计器的偏差在$p\sim n$处开始增长;越过峰值后,方差如回归中那样衰减,而偏差持续增长,两者最终都饱和在较大值。由于扩散模型训练时$m\gg1$,峰值被推至非常大的模型规模,因此模型位于峰值之前的上升分支上,此时恶性过拟合已经发挥作用。尽管如此,当与正则化结合时,过参数化仍然是有益的:在随机特征理论和U-Net实验中,最优正则化的大模型——分别通过岭惩罚或早停——优于任何未正则化的模型。

英文摘要

Conventional wisdom in deep learning holds that overparameterization---having more parameters $p$ than training samples $n$---is benign: larger models generalize better and, even without regularization, interpolating models generalize well, the test error following a double-descent curve. One might expect the same benign overfitting for diffusion models, whose training reduces to regression, i.e. to minimizing a quadratic score-matching loss. Yet the opposite is observed: overfitting here is catastrophic, driving the model into a memorization regime. We resolve this paradox by combining experiments on U-Nets trained on CelebA with a random-features model for which we derive closed-form learning curves. We show that with a fixed number $m$ of noise realizations per training sample, an interpolation peak does occur, but at $p\sim nm$ rather than at $p\sim n$ as in standard regression. The rise of the test loss, however, sets in much earlier, at $p\sim n$, independently of $m$. This overfitting is malign because, although the implicit regularization of training is fully at work, it drives the model toward the empirical score, which memorizes the training set, rather than toward the true score. A bias-variance decomposition pinpoints the mechanism: the bias of the score estimator starts to grow at $p\sim n$; past the peak the variance decays, as in regression, whereas the bias keeps growing and both saturate at a large value. Since diffusion models are trained with $m\gg1$, the peak is pushed to very large model sizes, and therefore sit on the rising branch that precedes it, where malign overfitting is already in play. Nevertheless, overparameterization remains beneficial when paired with regularization: in the random-features theory and in U-Net experiments, optimally regularized large models---via a ridge penalty or early stopping, respectively---outperform any unregularized models.

Comments44 pages, 17 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑