arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ACID:用于视频生成的自适应缓存

One Threshold Is Not Enough: Prompt-Invariant Caching Schedules for Video Diffusion

Om Agrawal, Saurabh Agarwal, Aditya Akella

arXiv 2607.12358首次发表:更新:

发表机构

UT Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视频扩散模型推理慢问题,提出ACID自适应缓存方法,通过监测漂移信号变化动态切换阈值,无需重新训练,可插入现有缓存方法,实验证明其能显著提升推理速度且质量退化可忽略。

AI 中文摘要

视频扩散模型能生成高质量内容,但由于其顺序去噪过程,推理速度较慢。基于缓存的加速方法通过重用中间模型输出解决此问题,如TeaCache、EasyCache和DiCache等动态方法,当累积漂移低于固定阈值τ时跳过昂贵的模型评估。我们发现这种权衡并非根本,存在关键步骤,在这些步骤选择性应用低阈值并在其他地方积极缓存,能在更高推理速度下恢复大部分保守缓存的质量。基于此,我们提出ACID,它能动态切换阈值,无需重新训练,可直接插入现有动态缓存方法。实验表明,ACID在视觉质量与推理速度上超越固定阈值方法,如在TeaCache和HunyuanVideo上,相比无缓存基线加速达2.16倍,相比保守固定阈值基线在质量退化可忽略时额外加速达38%。

英文摘要

Video diffusion models remain expensive because denoising repeatedly evaluates a costly model backbone. Dynamic caching methods like TeaCache, EasyCache, and DiCache reduce this cost by reusing prior outputs while an accumulated drift signal remains below a threshold $τ$. Although existing methods improve the drift signal, they hold $τ$ fixed throughout denoising. We show that fixed thresholds are structurally suboptimal: varying $τ$ across denoising achieves quality-latency tradeoffs unattainable by any fixed threshold across three caching methods with distinct drift signals. We further show that, surprisingly, each method's drift trajectory is nearly prompt-invariant across seven method-model combinations. The shape of the drift trajectory is determined by the model and caching method; therefore, a threshold schedule can be calibrated offline and reused across generations. Based on these findings, we present ACID (Adaptive Caching for vIDeo generation), which uses a low threshold during critical regions where drift changes rapidly and a high threshold across stable regions. ACID identifies these regions offline from the drift signal's second derivative, requires no training, and adds negligible runtime overhead. Across TeaCache, EasyCache, and DiCache on HunyuanVideo, Wan 2.1, and CogVideoX, ACID pushes the speed-quality Pareto frontier beyond fixed thresholds. On TeaCache with HunyuanVideo, it achieves 2.16x speedup over no caching and 38% additional speedup over a conservative fixed threshold, with less than 0.3 dB PSNR, 0.01 SSIM, and 0.01 LPIPS degradation relative to that fixed-threshold configuration.

Comments21 pages, 9 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑