arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从局部失配到全局影响:优化缓存重用策略以实现高效扩散模型

From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion

Xichen Ye, Yifan Wu, Zhikang Xie, Xiangyu Yue, Cheng Jin, Weizhong Zhang

arXiv 2608.13043首次发表:更新:

发表机构

Faculty of Engineering, The Chinese University of Hong Kong; School of Data Science, Fudan University(香港中文大学工程学院; 复旦大学数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对扩散模型缓存策略与生成质量失配问题,提出GCache双层优化框架,在Wan2.1模型上实现2.17倍加速且LPIPS降至0.0316,性能优于现有缓存策略。

AI 中文摘要

扩散模型在视觉生成领域已取得主导性能,但存在显著的推理开销问题。尽管基于缓存的加速方法已成为有前景的解决方案,但现有策略依赖局部相似性启发式规则,我们发现其与最终生成质量存在显著偏差。该偏差源于去噪轨迹上误差的非均匀传播与累积。为解决此问题,我们提出全局影响缓存(Global-Impact Cache,GCache)。首先,我们对误差传播上界建立严格的理论表征;考虑到该上界对于复杂、高度非凸的扩散模型可能过于保守,我们进一步用 Bernstein 形式重新参数化传播指数,并将缓存策略搜索重新表述为双层优化问题。具体而言,GCache 在内部目标中识别最优重用策略,同时在外部目标中将误差加权函数与生成质量损失对齐。该框架有效协调了理论严谨性与经验性能,学习在对视觉保真度影响最大的计算区域进行优先处理。大量实验表明,GCache 在图像和视频生成任务上均始终优于现有缓存策略。值得注意的是,在最先进的 Wan2.1 视频扩散模型上,GCache 保持了 2.17 倍的加速,同时显著提升了生成质量,将 LPIPS 从 0.1095 降至 0.0316。

英文摘要

Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which we identify as being significantly misaligned with final generation quality. This discrepancy stems from the non-uniform propagation and accumulation of errors along the denoising trajectory. To address this, we propose Global-Impact Cache (GCache). We first establish a rigorous theoretical characterization of the error propagation upper bound. Recognizing that this bound can be overly conservative for complex, highly non-convex diffusion models, we further reparameterize the propagation exponent with a Bernstein form and reformulate cache policy search as a bilevel optimization problem. In detail, GCache identifies an optimal reuse policy in the inner objective while aligning the error-weighting function with generation quality loss in the outer objective. This framework effectively reconciles theoretical rigor with empirical performance, learning to prioritize computation where it most impacts visual fidelity. Extensive experiments demonstrate that GCache consistently outperforms prior caching strategies on both video and image generation. Notably, on the state-of-the-art Wan2.1 video diffusion model, GCache maintains a 2.17x speedup while significantly enhancing generation quality, reducing LPIPS from 0.1095 to 0.0316.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑