AI 中文总结
研究提出DiFA框架,将扩散模型推理时的数据预测细化转为顺序状态估计问题,受卡尔曼滤波启发建立前向对齐时间一致性,引入偏差引导机制,实验表明该方法在多个指标上显著提升CIFAR-10和ImageNet的生成保真度。
AI 中文摘要
扩散模型的主流推理框架从根本上将生成视为数值积分问题。这种观点将模型视为精确估计器,忽略了去噪过程中固有的统计不确定性。在这项工作中,我们提出了前向过程对齐扩散预测(DiFA),这是一个无需训练的框架,将推理时间的数据预测细化重新构建为顺序状态估计问题。DiFA不是仅为数值积分而重用过去的输出,而是将沿反向轨迹的迭代数据预测视为相关观测,以建立前向对齐的时间一致性。受卡尔曼滤波启发,这种一致性根据结构一致性和噪声水平兼容性聚合历史预测。为了抵消时间一致性的过度平滑趋势,我们引入了偏差引导机制来自适应地保留残差细节。通过实验,DiFA在CIFAR-10和ImageNet上的包括FID、IS和FD-DINOv2等评估指标上取得了显著改进,表明将推理与前向统计结构对齐可显著提高生成保真度。
英文摘要
The prevailing inference framework for diffusion models formulates generation fundamentally as a problem of numerical integration. This perspective casts the model as an exact estimator, neglecting the inherent statistical uncertainty of the denoising process. In this work, we propose Forward-Process Aligned Diffusion prediction (\textbf{DiFA}), a training-free framework that reframes inference-time data prediction refinement as a sequential state estimation problem. Rather than reusing past outputs solely for numerical integration, DiFA treats iterative data predictions along the reverse trajectory as correlated observations to build a forward-aligned temporal consensus. Inspired by Kalman filtering, this consensus aggregates historical predictions according to structural consistency and noise-level compatibility. To counteract the over-smoothing tendency of temporal consensus, we introduce a deviation guidance mechanism to adaptively preserve residual details. Empirically, DiFA yields significant improvements on CIFAR-10 and ImageNet across the evaluated metrics, including FID, IS, and FD-DINOv2, demonstrating that aligning inference with the forward statistical structure substantially improves generative fidelity.
CommentsAccepted to ICML 2026. 23 pages, 7 figures