arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.21896cs.CV

具有显著性估计的过去和未来信息KV缓存策略在自回归视频扩散中

Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion

发表机构杭州理工大学,西安电子科技大学,中国杭州 · 西安电子科技大学,中国西安 · 河海大学,中国南京
查看机构详情
  • Hangzhou Institute of Technology, Xidian University, Hangzhou, China(杭州理工大学,西安电子科技大学,中国杭州)
  • Xidian University, Xi'an, China(西安电子科技大学,中国西安)
  • Hohai University, Nanjing, China(河海大学,中国南京)

机构由 AI 辅助整理,请以论文原文为准。

Hanmo Chen, Chenghao Xu, Xu Yang, Xuan Chen, Cheng Deng

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出PaFu-KV策略,通过显著性估计优化KV缓存,提升视频生成的质量与效率。

中文摘要 AI 辅助

视频生成对于数字媒体创作至关重要,最近在自回归视频生成方面的进展显著提高了实时视频合成的效率。然而,现有方法通常依赖于启发式KV缓存策略,这些策略忽略了长期视频生成中token重要性的差异。这导致了关键时空信息的丢失和冗余无效缓存的积累,从而降低了视频生成的质量和效率。为了解决这一限制,我们首先观察到token对视频生成的贡献具有高度的时间异质性,并相应地提出了一种新的过去和未来信息KV缓存策略(PaFu-KV)。具体来说,PaFu-KV引入了一个从双向教师模型中蒸馏而来的轻量级显著性估计头,用于估计显著性分数,使KV缓存能够保留信息性token并丢弃不相关的内容。该策略通过缩小KV缓存容量和减少推理时的内存占用,实现了更好的质量-效率权衡。在基准测试中的广泛实验表明,我们的方法在保持高质量视频生成的同时,实现了加速推理,从而实现了更高效的长视界视频生成。我们的代码将在论文接受后发布。

英文摘要

Video generation is pivotal to digital media creation, and recent advances in autoregressive video generation have markedly enhanced the efficiency of real-time video synthesis. However, existing approaches generally rely on heuristic KV Cache policies, which ignore differences in token importance in long-term video generation. This leads to the loss of critical spatiotemporal information and the accumulation of redundant, invalid cache, thereby degrading video generation quality and efficiency. To address this limitation, we first observe that token contributions to video generation are highly time-heterogeneous and accordingly propose a novel Past- and Future-Informed KV Cache Policy (PaFu-KV). Specifically, PaFu-KV introduces a lightweight Salience Estimation Head distilled from a bidirectional teacher to estimate salience scores, allowing the KV cache to retain informative tokens while discarding less relevant ones. This policy yields a better quality-efficiency trade-off by shrinking KV cache capacity and reducing memory footprint at inference time. Extensive experiments on benchmarks demonstrate that our method preserves high-fidelity video generation quality while enables accelerated inference, thereby enabling more efficient long-horizon video generation. Our code will be released upon paper acceptance.

↑