发表机构
Nanyang Technological University; Tsinghua University; Princeton University(南洋理工大学; 清华大学; 普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DeCoPrune利用去噪一致性信号剪枝KV缓存,在自回归视频扩散中实现高效长程信息保留,显著降低推理成本并提升速度。
AI 中文摘要
自回归视频扩散支持流式生成和交互式控制,但其KV缓存会随着生成的历史内容而增长。现有的压缩策略通过固定窗口丢弃历史信息,或通过局部注意力和相似性信号选择令牌,而没有直接衡量某个块是否贡献了超出保留上下文的信息。我们提出了DeCoPrune,一种无需训练的方法,将缓存压缩视为去噪一致性问题。我们发现,中间干净预测与最终去噪值之间差异较大的令牌,往往携带较难从保留上下文中预测的视觉证据。DeCoPrune利用这一模型内在信号,在长期缓存中保留高差异令牌,同时剪枝低差异令牌。为了评估信息保留能力,我们引入了CMBench,包含58个约一分钟的生成或真实世界上下文片段,以及116个需要回忆早期事件或物体的“重现”或“重访”续写任务。使用LingBot World v2进行的实验表明,DeCoPrune在0-1尺度上达到了0.6701的DINO分数,累计历史KV令牌数量减少了85.43%,续写生成速度比FullKV提升了4.14倍。其头部特化变体在86.19%的剪枝率下达到了0.6783,接近FullKV的0.6803分数,并在相似预算下超过了所评估的压缩基线。这些结果表明,去噪一致性可以在降低自回归推理成本的同时支持长距离信息保留。我们的项目主页为https URL。代码可在https URL获取,基准测试可在https URL获取。
英文摘要
Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows continuously with the generated history. Existing compression strategies either discard history using fixed windows or select tokens through local attention and similarity signals, which do not directly measure whether the current chunk contributes information beyond the retained context. We introduce DeCoPrune, a training-free method that treats cache compression as a denoising-consistency problem. We find empirically that denoising difficulty provides a useful proxy for a token's value in long-term retention: tokens with larger step-to-final discrepancies tend to carry visual evidence that is less predictable from the retained context. DeCoPrune measures each current-chunk token's denoising difficulty using the discrepancy between its intermediate clean prediction and final denoised value, retaining high-discrepancy tokens in the long-term cache while pruning those with low discrepancy. To evaluate information retention, we introduce CMBench, comprising 58 approximately one-minute generated or real-world context episodes and 116 Reappear or Revisit continuation tasks that require recalling specific previously observed objects or scenes. Experiments with LingBot World v2 show that DeCoPrune preserves near-FullKV long-range recall while pruning over 85% of historical KV tokens and accelerating continuation generation by over $4\times$, substantially outperforming the evaluated compression baselines at comparable budgets. These results indicate that denoising consistency can serve as a model-intrinsic signal for retaining long-range information while reducing autoregressive inference cost. Our project homepage is https://decoprune.github.io. The code is available at https://github.com/DeCoPrune/CMBench, and the benchmark at https://huggingface.co/datasets/Aoraku/CMBench.