arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩散模型中的伪随机流是影响生成质量的可学习输入

Noise in Diffusion Models Is a Learnable Input

Shengzhi Deng, Chenqi Ye, Yanze Guo

arXiv 2608.02575首次发表:更新:

发表机构

School of Mathematics, Harbin Institute of Technology; Software College, Northeastern University(哈尔滨工业大学数学学院; 东北大学软件学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现扩散模型的伪随机流是可学习输入,会影响生成质量,相关指标与生成退化相关,伪随机源兼具分布选择与结构化输入属性。

AI 中文摘要

扩散模型依赖随机输入,但在有限精度硬件上,其消耗的“随机性”表现为伪随机规则生成的确定性数值轨道。可访问的轨道结构可成为可学习输入,影响训练与生成,因为实际损失及其梯度取决于每次优化步骤消耗的具体伪随机值。小型多层感知器从轨道近期历史预测其下一个值,衡量序列的可预测性;扩散探针则在保留扩散架构与训练目标的同时,将真实图像替换为在线随机张量,衡量目标系统能否利用轨道结构。在控制边缘统计并排除明显的动力学及有限精度失效后,剩余轨道在MNIST和CIFAR-10数据集上仍产生显著不同的扩散损失与生成质量,且这两项指标与宏观生成退化呈强秩相关,尽管局部排名存在差异。经独立同分布(IID)基线归一化后,探针损失与真实数据扩散损失近似遵循经验幂律,且在两个数据集上指数不同。这些结果表明,伪随机源不仅是分布选择,还是依赖模型的结构化输入。

英文摘要

Stochastic learning objectives are typically written as expectations over abstract random variables. Actual training, however, uses concrete random inputs that enter both the realized loss and its gradient. Structure in these inputs that is accessible to the learning system can therefore be learned and exploited. Much prior structured-noise work asks how noise should be distributed or designed; we instead ask what structure in the concrete realized randomness becomes exploitable by the learner. We develop this general view and analyze its mechanism in diffusion models: in noise prediction, clean data and realized noise jointly form the noisy input, so the model can improve prediction by learning clean-data regularities or exploiting structure introduced through the noise, and the two routes can interact. Using pseudorandom streams as controlled, reproducible instances of structured noise, we provide mechanistic evidence on MNIST and CIFAR-10: random-role ablations localize the dominant effect to diffusion noise; in a diffusion probe, structured-noise training can reduce prediction loss below the IID reference, but replacing the test noise with IID reverses this advantage; and shuffling the same values largely removes the source-dependent loss reduction. This learned dependence can also affect generation. The same view offers a unified interpretation of data-dependent noise assignment, noise-based backdoors, and temporally correlated noise in video diffusion: although these methods introduce different structures, all alter what the model can exploit through noise and its interaction with clean-data learning. Our results indicate that diffusion noise is not merely a passive stochastic perturbation, but a learnable---and therefore potentially designable---input dimension.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑