arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于得分的生成模型中随机梯度下降的非渐近收敛性

Non-asymptotic Convergence of Stochastic Gradient Descent in Score-based Generative Models

Stanislas Strasman, Sobihan Surendran, Sylvain Le Corff

arXiv 2607.04775首次发表:更新:

发表机构

Sorbonne Université and Université Paris Cité, CNRS, LPSM; LOPF, Califrais’ Machine Learning Lab(索邦大学和巴黎cité大学,法国国家科学研究中心,巴黎高等师范学院; LOPF,Califrais机器学习实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对基于得分的生成模型训练优化性质不明问题,本文对两类场景下的随机梯度下降开展理论分析,推导收敛速率与误差边界,为实际权重选择提供理论指导。

AI 中文摘要

基于得分的生成模型(SGMs)已在各类应用的数据生成任务中取得优异表现。尽管其采样过程的统计特性已得到越来越充分的研究,但其训练过程背后的优化动态仍探索不足。SGMs通常通过最小化加权去噪得分匹配目标完成训练,目前针对随机梯度的优化保证仍较为有限。本工作研究面向SGMs的随机梯度下降(SGD),在两个互补场景下给出相关结论:第一,针对通用的得分参数化形式,我们建立了SGD在加权去噪得分匹配目标上的非凸收敛速率,该速率显式依赖于与调度相关的加权因子;第二,针对过参数化两层ReLU网络,我们提出了适配随机梯度扩散训练的神经正切核分析方法,推导得到SGD轨迹上的得分近似误差边界。最终,本文的分析量化了重加权因子在得分近似误差中的作用,为实际应用中的权重选择提供了理论指导。

英文摘要

Score-based Generative Models (SGMs) have achieved impressive performance in data generation across a wide range of applications. While the statistical properties of their sampling procedures are increasingly well understood, the optimization dynamics underlying their training remain less explored. SGMs are typically trained by minimizing a weighted denoising score-matching objective, yet optimization guarantees with stochastic gradients remain limited. In this work, we study Stochastic Gradient Descent (SGD) for SGMs, contributing results in two complementary regimes. For general score parameterizations, we derive a non-convex analysis of SGD for the weighted denoising score-matching objective, making explicit how the resulting optimization bound depends on the loss weighting and time-sampling distribution. We then consider overparameterized two-layer ReLU networks and develop a Neural Tangent Kernel analysis tailored to diffusion training with stochastic gradients, yielding score-approximation error bounds along the SGD trajectory. Our analysis quantifies the role of the reweighting factor in these bounds, providing a theoretical characterization of weighting choices used in practice.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑