arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2512.12604cs.CV

无空闲缓存:通过极端瘦身缓存加速扩散模型

No Cache Left Idle: Accelerating diffusion model via Extreme-slimming Caching

  • Tsinghua University(清华大学)
  • Central Media Technology Institute, Huawei(华为中央媒体技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

Tingyan Wen, Haoyu Li, Yihuang Chen, Xing Zhou, Lifei Zhu, Xueqian Wang

更新

AI总结:

X-Slim通过双阈值缓存策略,高效利用时间步、块和令牌层面的冗余,实现扩散模型加速,减少延迟并提升生成质量。

AI中文摘要:

扩散模型实现了卓越的生成质量,但计算开销随步骤数、模型深度和序列长度而增加。特征缓存有效,因为相邻时间步具有高度相似的特征。然而,存在一个固有的权衡:激进的时间步重用能提供大的加速,但容易跨临界线,损害保真度,而块级或令牌级重用更安全但计算节省有限。我们提出了X-Slim(eXtreme-Slimming Caching),一种训练无关、基于缓存的加速器,据我们所知,是第一个统一框架,利用时间步、结构(块)和空间(令牌)层面的可缓存冗余。X-Slim 不只是混合级别,而是引入了双阈值控制器,将缓存转化为“推-然后-润色”过程:它首先在时间步层面将重用推至早期预警线,然后切换到轻量级块级和令牌级刷新来润色剩余的冗余,并在跨过临界线时触发完整推断以重置累积误差。在每个层面,上下文感知指标决定何时以及在哪里缓存。在多样化的任务中,X-Slim 推动了速度-质量前沿。在FLUX.1-dev和HunyuanVideo上,它将延迟减少了高达4.97倍和3.52倍,同时感知损失极小。在DiT-XL/2上,它实现了3.13倍的加速,并在先前方法上提高了2.42的FID。

英文摘要:

Diffusion models achieve remarkable generative quality, but computational overhead scales with step count, model depth, and sequence length. Feature caching is effective since adjacent timesteps yield highly similar features. However, an inherent trade-off remains: aggressive timestep reuse offers large speedups but can easily cross the critical line, hurting fidelity, while block- or token-level reuse is safer but yields limited computational savings. We present X-Slim (eXtreme-Slimming Caching), a training-free, cache-based accelerator that, to our knowledge, is the first unified framework to exploit cacheable redundancy across timesteps, structure (blocks), and space (tokens). Rather than simply mixing levels, X-Slim introduces a dual-threshold controller that turns caching into a push-then-polish process: it first pushes reuse at the timestep level up to an early-warning line, then switches to lightweight block- and token-level refresh to polish the remaining redundancy, and triggers full inference once the critical line is crossed to reset accumulated error. At each level, context-aware indicators decide when and where to cache. Across diverse tasks, X-Slim advances the speed-quality frontier. On FLUX.1-dev and HunyuanVideo, it reduces latency by up to 4.97x and 3.52x with minimal perceptual loss. On DiT-XL/2, it reaches 3.13x acceleration and improves FID by 2.42 over prior methods.

补充信息

↑