arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Koopman观测器用于扩散加速:利用浅层测量修正特征预测

Koopman Observers for Diffusion Acceleration: Correcting Feature Forecasts with Shallow Measurements

Hanru Bai, Yuanchao Xu, Fengyi Li

arXiv 2610.10366首次发表:更新:

发表机构

ETH Zurich; Max Planck Institute for Intelligent Systems; Kyoto University; MIT(苏黎世联邦理工学院; 马克斯·普朗克智能系统研究所; 京都大学; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出观测修正的Koopman框架,利用浅层特征修正深层特征预测,加速扩散采样,在CIFAR-10和ImageNet子集上降低特征MSE并实现近两倍加速。

AI 中文摘要

特征缓存通过用先前计算的激活值的预测替换昂贵的网络评估来加速扩散采样。然而,仅基于过去特征的预测无法直接纳入当前去噪状态的变化。我们研究是否可以利用廉价的、新计算的特征作为观测来修正这些预测。我们引入了一个观测修正的Koopman框架,用于加速冻结的扩散模型。利用校准轨迹,我们识别出有限维、时间依赖的Koopman近似,共同描述浅层和深层网络特征的增量。在加速采样过程中,这些算子预测昂贵的深层特征的演化,而观测到的浅层特征中的创新修正预测状态。周期性完整评估刷新观测器,所有生成模型参数保持不变。该公式使得时间预测和观测修正的受控比较成为可能。在每数据集三次10,000图像运行中,相对于相同四部分步调度下的通道仿射预测,我们的方法在CIFAR-10上将配对Inception特征MSE降低了19.9%,在十类ImageNet子集上降低了11.9%。匹配的消融研究将额外4.54%和4.67%的降低归因于观测修正。该观测器相对于DDIM-50实现了1.89倍和1.85倍的实测加速,支持在不重新训练去噪器的情况下提高参考采样器保真度。

英文摘要

Feature caching accelerates diffusion sampling by replacing expensive network evaluations with predictions from previously computed activations. However, forecasts based only on past features cannot directly incorporate changes in the current denoising state. We investigate whether inexpensive, freshly computed features can serve as observations for correcting these predictions. We introduce an observation-corrected Koopman framework for accelerating frozen diffusion models. Using calibration trajectories, we identify finite-dimensional, time-dependent Koopman approximations that jointly describe the increments of shallow and deep network features. During accelerated sampling, these operators predict the evolution of expensive deep features, while innovations in the observed shallow features correct the predicted state. Periodic full evaluations refresh the observer, and all generative-model parameters remain unchanged. This formulation enables controlled comparisons of temporal prediction and observation correction. Across three 10,000-image runs per dataset, our method reduces paired Inception-feature MSE by $19.9\%$ on CIFAR-10 and $11.9\%$ on a ten-class ImageNet subset relative to channelwise affine prediction under the same four-partial-step schedule. Matched ablations attribute additional reductions of $4.54\%$ and $4.67\%$ to observation correction. The observer achieves $1.89\times$ and $1.85\times$ measured speedups over DDIM-50, supporting improved reference-sampler fidelity without retraining the denoiser.

Comments14 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑