从得分学习到离散采样:扩散模型的端到端泛化分析
From Score Learning to Discretized Sampling: An End-to-End Generalization Analysis of Diffusion Models
- School of Mathematical Sciences, Nankai University(南开大学数学科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究基于实际残差网络架构的得分扩散模型,建立统一收敛和泛化框架,分析从得分函数学习到采样过程的特性,将生成误差分解为四个分量,定量刻画训练样本大小等因素对扩散模型样本保真度的控制。
AI中文摘要:
尽管基于得分的扩散模型在经验上取得了成功,但对于有限样本学习、网络参数化和数值离散化如何共同决定生成质量,仍缺乏完整的理论理解。现有采样分析往往基于神谕得分或预先指定的误差阈值来评估生成性能。在这项工作中,我们为基于实际残差网络类型架构参数化的得分扩散模型建立了一个统一的收敛和泛化框架。我们分析了从得分函数的实际有限样本、离散时间学习问题到理想连续时间、总体水平目标的泛化和收敛特性。基于得分函数学习问题的泛化结果,我们分析了由学习到的得分函数诱导的采样过程,并为生成的终端分布提供了端到端的总变差距离估计。该估计将整体生成误差明确分解为四个可解释的分量:正向过程的截断误差、反向时间离散化误差、包含有限数据和正向时间离散化的泛化误差以及训练优化差距。我们的结果定量地刻画了训练样本大小、时间离散化网格和优化精度如何共同控制扩散模型生成样本的最终保真度。
英文摘要:
Despite the empirical success of score-based diffusion models, a complete theoretical understanding of how finite-sample learning, network parameterization, and numerical discretization jointly dictate generative quality remains underdeveloped. Existing sampling analyses often evaluate the generative performance conditional on an oracle score or a pre-specified error threshold. In this work, we establish a unified convergence and generalization framework for score-based diffusion models parameterized by practical ResNet-type architectures. We analyze the generalization and convergence properties from the practical finite-sample, discrete-time learning problem of the score function to the ideal continuous-time, population-level objective. Based on the generalization result of the learning problem of score function, we analyze the sampling process induced by the learned score function and provide an end-to-end total variation distance estimate for the generated terminal distribution. This estimate explicitly decomposes the overall generative error into four interpretable components: the truncation error of the forward process, the reverse-time discretization error, the generalization error incorporating both finite data and forward-time discretization, and the training optimization gap. Our results quantitatively characterize how the training sample size, temporal discretization grids, and optimization accuracy jointly control the final fidelity of samples generated by diffusion models.