arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OnlineCache:用于高效扩散推理的带误差校正的动态缓存策略学习

OnlineCache: Learning Dynamic Caching Policies with Error Correction for Efficient Diffusion Inference

Zhikang Xie, Xichen Ye, Yifan Wu, Haoshen Yu, Li chenan, Peizhu Gong, Weizhong Zhang, Cheng Jin

arXiv 2607.29398首次发表:更新:

发表机构

College of Computer Science and Artificial Intelligence, Fudan University; School of Data Science, Fudan University(复旦大学计算机科学与人工智能学院; 复旦大学数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对扩散模型推理延迟高的问题,提出动态缓存框架OnlineCache,结合策略梯度与可学习校正器,实现样本与时间步的自适应资源分配,在FLUX.1-dev等模型上实现加速且优于现有基线。

AI 中文摘要

扩散模型彻底改变了生成式任务,但由于迭代去噪会带来高延迟。基于缓存的策略通过复用中间特征来加速推理,但它们在很大程度上依赖于静态、与样本无关的调度。本文实证验证了这种僵化忽略了两个事实:(i)不同提示的生成难度各不相同,需要自适应资源分配——复杂输入需要更多计算,而更简单的输入需要更少;(ii)不同时间步的误差敏感性会波动,静态策略可能会缓存高误差步骤,或在低误差步骤上浪费计算。因此,我们提出OnlineCache,这是一个动态缓存框架,联合学习何时缓存以及如何校正近似误差。我们利用策略梯度训练一个轻量网络,以实现自适应速度-质量权衡,并融入一个可学习校正器来缓解缓存引发的误差。两个模块在双层优化框架下联合优化,策略针对全局生成质量,校正器则最小化局部误差。我们的方法会在样本和时间步之间自动分配计算资源,提升整体生成质量。大量实验表明其具有明显优势:在FLUX.1-dev模型上,OnlineCache实现了近3倍加速,同时保留了生成保真度;在DiT和CogVideoX上,它同样实现了有竞争力的加速且不损害质量;在所有场景中,它始终优于现有的基于缓存的加速基线。

英文摘要

Diffusion models have revolutionized generative tasks but incur high latency due to iterative denoising. While cache-based strategies accelerate inference by reusing intermediate features, they largely rely on static, sample-agnostic schedules. We argue that this rigidity overlooks two facts empirically validated in this paper: (i) generation difficulty varies across prompts, requiring adaptive resource allocation--complex inputs demand more computation while simpler ones require less; (ii) error sensitivity fluctuates across timesteps, where static policies may cache high-error steps or waste computation on low-error ones. We therefore propose OnlineCache, a dynamic caching framework that jointly learns when to cache and how to correct approximation errors. We leverage policy gradient to train a lightweight network for adaptive speed-quality trade-offs, and incorporate a learnable corrector to mitigate caching-induced errors. Both modules are jointly optimized under a bilevel optimization framework, with the policy targeting global generation quality and the corrector minimizing local errors. Our method automatically allocates computational resources across both samples and timesteps, improving overall generation quality. Extensive experiments demonstrate clear superiority. On FLUX.1-dev model, OnlineCache achieves nearly 3 speedup while preserving generation fidelity. On DiT and CogVideoX, it similarly delivers competitive acceleration without compromising quality; across all scenarios, it consistently outperforms existing cache-based acceleration baselines.

CommentsDynamic timestep-level cache method for diffusion acceleration via policy gradient

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑