AI 中文总结
针对扩散与流模型的计算及重建瓶颈,提出Interval Denoiser框架,经训练在ImageNet 256x256上实现少步生成的低FID值,提升了少步生成性能。
AI 中文摘要
现代扩散模型和基于流的模型正越来越多地转向少步、无隐变量生成,以规避多步采样的计算开销和外部自编码器的重建瓶颈。我们提出了Interval Denoiser(区间去噪器),这是一个用于无隐变量生成的理论严谨框架,它直接从流匹配常微分方程(flow matching ODE)推导而来,为中间轨迹状态建立了精确的解析映射。与现有公式不同,我们的预测结果被证明在任意时间区间内都位于低维流形上,这使得直接在像素上运行的网络能够轻松进行回归。此外,通过避免经验代数替换,我们的公式能正确分离纯时间导数,以防止有偏梯度评估并确保精确的一阶优化。通过分析该目标,我们为框架配备了残差裁剪和时间采样课程,从而实现了有效的长区间训练并提升了少步性能。在ImageNet 256x256上从头开始训练后,我们的模型在1步(1-NFE)时FID达到4.55,在2步(2-NFE)时FID达到3.98,且无需感知损失。
英文摘要
Modern diffusion and flow-based models are increasingly moving toward few-step, latent-free generation to bypass the computational overhead of multi-step sampling and the reconstruction bottlenecks of external autoencoders. We propose the Interval Denoiser, a theoretically rigorous framework for latent-free generation. Derived directly from the flow matching ODE, it establishes an exact analytical mapping for intermediate trajectory states. Unlike prior formulations, our prediction is shown to reside on a low-dimensional manifold across any time interval, making the regression tractable for a network operating directly on pixels. Furthermore, by avoiding empirical algebraic substitutions, our formulation correctly isolates the pure time derivative to prevent biased gradient evaluations and ensure exact first-order optimization. By analyzing this objective, we equip our framework with residual clipping and a time-sampling curriculum, enabling effective long-interval training and improving few-step performance. Trained from scratch on ImageNet 256x256, our model achieves an FID of 4.55 in one step (1-NFE) and 3.98 in two steps (2-NFE) without perceptual losses.