发表机构
Yale University; Institute for Foundations of Data Science, Yale University(耶鲁大学; 耶鲁大学数据科学基础研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对惰性高维 regime 中训练的扩散模型,推导了梯度流训练下的精确风险轨迹,揭示了其泛化、记忆与过拟合的阶段机制,为生成式建模提供了理论分析框架。
AI 中文摘要
现代基于分数的生成模型在图像、音频和视频合成等高维任务中取得了显著的经验成功。这些模型将分布学习简化为一系列回归问题,若在有限数据上精确求解,最终会重现训练样本。因此,它们的泛化能力必然源于训练过程中隐式或显式的正则化。在本研究中,我们针对监督惰性训练 regime 中超参数化神经网络的良性过拟合与算法正则化理论,构建了生成式对应理论。我们研究了带内积核的向量值再生核希尔伯特空间中的去噪分数匹配。在比例高维 regime $n\backsim d$ 中,我们推导了梯度流训练下的精确风险轨迹。这些轨迹呈现出三个由定性不同的估计器主导的阶段:可泛化的谱估计器、对训练目标进行插值的具有局部峰值的纯噪声分数,以及对数据进行记忆的经验贝叶斯估计器。随后,我们分析了这些估计器如何沿反向时间 SDE 组合,并表征了所得样本的分布。该分析揭示了监督学习中熟悉的机制,包括核线性化和核非线性部分的自诱导正则化,但也揭示了生成式建模特有的独特现象。
英文摘要
Modern score-based generative models have achieved remarkable empirical success in high-dimensional tasks such as image, audio, and video synthesis. These models reduce distribution learning to a sequence of regression problems that, if solved exactly on finite data, would ultimately reproduce the training samples. Their ability to generalize must therefore arise from the implicit or explicit regularization during training. In this work, we develop a generative counterpart to the theory of benign overfitting and algorithmic regularization for overparameterized neural networks in the supervised lazy-training regime. We study denoising score matching in a vector-valued reproducing kernel Hilbert space with an inner-product kernel. In the proportional high-dimensional regime $n\asymp d$, we derive exact risk trajectories under gradient flow training. These trajectories exhibit three phases governed by qualitatively distinct estimators: a spectral estimator that generalizes, a pure-noise score with localized peaks that interpolate the training objective, and an empirical Bayes estimator that memorizes the data. We then analyze how these estimators combine along the reverse-time SDE and characterize the distribution of the resulting samples. The analysis reveals familiar mechanisms from supervised learning, including kernel linearization and self-induced regularization from the nonlinear part of the kernel, but also reveals a distinct phenomenology specific to generative modeling.
Comments100 pages, 3 figures