发表机构
Great Wall Motors Co. Ltd.; China Patent Information Center(长城汽车股份有限公司; 中国专利信息中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推导出去噪分数匹配的损失下限的精确表达式,关联薛定谔桥与费希尔几何,揭示扩散模型训练损失的固有信息几何成分,指出不同噪声设置下原始损失无法一致排序的原因。
AI 中文摘要
去噪分数匹配通过对条件分数进行回归来训练扩散模型,尽管生成过程最终需要边缘分数。这两个目标共享相同的总体极小值点,但在固定噪声状态下,条件目标仍为随机量,会导致训练损失中出现不可约的超额损失。我们分离出该超额损失并证明,在温和正则性假设下,对于一般的 corruption 核,其恰好是条件端点族的 Fisher-Rao 度量的迹,沿扩散轨迹积分所得。这为去噪目标提供了精确的条件方差分解,并将扩散潜空间中观测到的信息几何确定为训练损失的固有组成部分。我们从薛定谔桥变分原理推导出该结果,其中理想目标表现为路径空间相对熵的超额项。对于 corruption 扩散,费希尔项与噪声状态丢失关于干净数据的互信息的速率成正比,将损失下限分为由数据决定的信息流和由 corruption schedule 及目标决定的权重。在高斯情况下,这给出了下限的闭式形式,恢复了连续时间目标的重参数化不变性,并将其高信噪比散度与数据的信息维数关联起来。最后,我们表明,使用不同噪声范围或权重得到的原始损失不一定能一致地对模型进行排序,因为它们包含不同的加性下限,并对比了训练所观测到的二阶几何与数值采样误差所涉及的三阶条件统计量。
英文摘要
Denoising score matching trains diffusion models by regressing onto a conditional score, although generation ultimately requires the marginal score. The two objectives share the same population minimizer, but the conditional target remains random at fixed noisy state and introduces an irreducible excess in the training loss. We isolate this excess and show that, for a general corruption kernel under mild regularity assumptions, it is exactly the trace of the Fisher--Rao metric of the conditional endpoint family, integrated along the diffusion trajectory. This gives an exact conditional-variance decomposition of the denoising objective and identifies the information geometry observed in diffusion latent spaces as an intrinsic component of the training loss. We derive the result from a Schr"odinger bridge variational principle, in which the ideal objective arises as excess path-space relative entropy. For corruption diffusions, the Fisher term is proportional to the rate at which the noisy state loses mutual information about the clean data, separating the loss floor into an information flow determined by the data and a weight determined by the corruption schedule and objective. In the Gaussian case, this yields a closed form for the floor, recovers reparametrization invariance of the continuous-time objective, and relates its high-SNR divergence to the information dimension of the data. Finally, we show that raw losses obtained with different noise ranges or weightings need not rank models consistently because they contain different additive floors, and contrast the second-order geometry seen by training with the third-order conditional statistics entering numerical sampling error.
Comments28 pages, 4 figures