发表机构
Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
一步生成模型将扩散多步去噪压缩为单次前向传播,发现去噪计算沿网络深度展开,且可压缩16.6倍参数,表明时间计算被重组为深度计算。
AI 中文摘要
最近的一步生成模型浪潮,通过蒸馏或学习到的流映射压缩扩散的多步轨迹,已达到一个临界点,能够生成高质量图像。在此,我们提出一个由这些进展自然引发的问题:当生成被压缩为单次前向传播时,多步扩散的去噪轨迹会发生什么?我们提供了一个称为“深度作为时间”的经验观察:多步扩散在采样步骤中执行的去噪计算似乎展开在单次前向传播的深度上,并且可以通过使用模型自身的输出头解码中间层来恢复。最有趣的是,我们表明这种逐深度计算取决于流映射被训练来解决的传输任务。最令人惊讶的情况是MeanFlow,其中探测较短的传输间隔揭示了在单次网络评估中的去噪和再噪声化。相比之下,没有时间索引传输任务训练的生成器,如漂移模型,不表现出相同的逐深度去噪。因此,我们表明表现出逐深度去噪现象的模型在逐层计算上更可压缩:一个MeanFlow SiT-L/2模型可以被压缩16.6倍参数到一个单一时间条件块中。我们为这种去噪然后再噪声化的行为提供了解释,并表明当我们明确地将逐层计算视为一个流时,一个单一时间条件块可以被训练为跨层去噪,将MeanFlow SiT-L/2模型压缩16.6倍参数。总之,这些结果表明,扩散的时间计算并未被一步生成消除,而是重新组织在网络深度上。
英文摘要
The recent wave of one-step generative models, which compress the multi-step trajectory of diffusion via either distillation or learned flow maps, has reached an inflection point where they can generate high-quality images. Here, we ask a natural question that follows from these advances: what happens to the denoising trajectory of multi-step diffusion when generation is compressed into a single forward pass? We offer an empirical observation we call \textit{depth as time}: the denoising computation that multi-step diffusion performs across sampling steps appears to unfold across the depth of a single forward pass, and can be recovered by decoding intermediate layers with the model's own output head. Most interestingly, we show that this depthwise computation depends on the transport task a flow map is trained to solve. The most surprising case is MeanFlow, where probing shorter transport intervals reveals both denoising and renoising within a single network evaluation. In contrast, generators trained without a time-indexed transport task, such as drifting models, do not exhibit the same depthwise denoising. Consequently, we show that models that exhibit the depthwise denoising phenomenon are more compressible across the layerwise computation: a MeanFlow \texttt{SiT-L/2} model can be compressed by $16.6\times$ in parameters into a single time-conditioned block. We offer an explanation for this denoise-then-renoise behavior and show that, when we treat the layerwise computation explicitly as a flow, a single time-conditioned block can be trained to denoise across layers, compressing a MeanFlow \texttt{SiT-L/2} model by $16.6\times$ in parameters. Together, these results suggest that the temporal computation of diffusion is not eliminated by one-step generation, but reorganized across network depth.