发表机构
École normale supérieure; Université PSL; CNRS; Sorbonne Université; Université Paris Cité(巴黎高等师范学院; 巴黎文理研究大学; 法国国家科学研究中心; 索邦大学; 巴黎西岱大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过高维分析证明,基于分数的生成模型中加权调度$w(t)$决定多模态数据特征(模式方向与权重)的学习顺序与时间尺度,并指出靠近物种形成时间的加权量对学习动力学至关重要。
AI 中文摘要
基于分数的生成模型通过积分一个将高斯噪声携带到目标分布上的时间相关漂移来生成新样本。在实践中,该漂移由神经网络建模,并在时间$t$上以加权调度$w(t)$对损失进行积分训练。沿着反向动力学,对于多模态分布,轨迹在狭窄的时间窗口内承诺到目标的模式,即\textit{物种形成时间}。在本工作中,聚焦于高维数据,我们将积分损失分解为其单时间贡献,并在固定信噪比$\Lambda(t)$下分析每个贡献:我们表明$\Lambda(t)$设定了多模态目标的每个特征——模式方向及其相对权重——在训练期间被获取的速率。关键的是,在高$\Lambda(t)$时,所有模式方向同时被获取,在一个对其幅度不敏感的单一时间尺度上,而相对权重则完全不被学习。仅在接近物种形成时间时,其中$\Lambda(t)$变为一阶量,所有特征才变得可学习,每个特征在其自身的时间尺度上:权重与方向同时获取,方向以由其相对幅度设定的速率获取。对于在时间积分目标上训练的模型,学习动力学随后由有效位于物种形成时间附近的加权量所支配,这为$w(t)$设计选择提供了见解。这些结果来自对不平衡和高斯混合模型训练动力学的精确高维分析。在图像和人类基因组单倍型生成上的数值实验在更复杂的设置中恢复了预测的学习时间尺度层次。
英文摘要
Score-based generative models generate new samples by integrating a time-dependent drift that carries Gaussian noise onto the target distribution. In practice this drift is modeled by a neural network, trained on a loss integrated over time $t$ with a weighting schedule $w(t)$. Along the backward dynamics, and for multi-modal distributions, trajectories commit to modes of the target within a narrow time window, the \textit{speciation time}. In this work, focusing on high-dimensional data, we decompose the integrated loss into its single-time contributions and analyze each at fixed signal-to-noise ratio $Λ(t)$: we show that $Λ(t)$ sets the rate at which each feature of a multimodal target - the mode directions and their relative weights - is acquired during training. Crucially, at high $Λ(t)$ all mode directions are acquired together, on a single timescale insensitive to their amplitudes, while the relative weights are not learned at all. Only near the speciation time, where $Λ(t)$ becomes of order one, do all features become learnable, each on its own timescale: the weights are acquired jointly with the directions, and the directions at rates set by their relative amplitudes. For models trained on time-integrated objectives, the learning dynamics is then governed by how much of the weighting effectively sits near the speciation time, which provides insights on $w(t)$ design choices. These results follow from an exact high-dimensional analysis of the training dynamics of unbalanced and hierarchical Gaussian mixtures. Numerical experiments on image and human genome haplotype generation recover the predicted hierarchy of learning timescales in more complex settings.
CommentsMain text : 9 pages / 5 figures Supplemental : 21 pages / 1 figure