AI 中文总结
研究为一类无记忆噪声过程的扩散模型开发基于GEXIT函数的守恒定律,统一离散和连续扩散似然性表征,利用局部性属性简化训练,通过在合成数据和标准基准上验证,揭示有限容量去噪器对不同噪声后验近似精度差异影响性能。
AI 中文摘要
自回归模型通过链式法则优化精确数据似然,而扩散模型通常用去噪目标训练。我们基于广义外部信息转移(GEXIT)函数为一类无记忆噪声过程开发守恒定律,表明数据-模型交叉熵(CE)可精确表征为沿噪声路径的局部信息论导数的积分。这为离散和连续扩散的似然性提供统一表征,高斯情况简化为著名的互信息-最小均方误差(I-MMSE)关系。直接结果是局部性属性,可仅用沿噪声路径的边际后验计算信息论导数。训练简化为通过最小化负对数似然学习边际后验。虽然守恒定律意味着熵不依赖于噪声路径,但有限容量去噪器对不同噪声类型的后验近似精度不同,导致性能差异。我们在合成马尔可夫源和标准基准(包括text8和CIFAR-10)上验证了这些预测。
英文摘要
While autoregressive models optimize the exact data likelihood via the chain rule, diffusion models are typically trained with denoising objectives. We develop conservation laws based on generalized extrinsic information transfer (GEXIT) functions for a broad class of memoryless noise processes, showing that the data--model cross-entropy (CE) can be characterized exactly as an integral of local information-theoretic derivatives along the noise path. This yields a unified characterization of the likelihood for discrete and continuous diffusion, with the Gaussian case reducing to the well-known mutual information--minimum mean-square error (I-MMSE) relationship. An immediate implication is a locality property: one can compute the information-theoretic derivatives using only the marginal posteriors along the noise path. As a result, training reduces to learning the marginal posteriors by minimizing the negative log-likelihood. While the conservation law implies that the entropy does not depend on the noise path, finite-capacity denoisers approximate the posteriors with varying accuracy across noise types, leading to differences in performance. We validate these predictions on synthetic Markov sources and standard benchmarks, including text8 and CIFAR-10.