发表机构
Harbin Institute of Technology (Shenzhen); National University of Singapore(哈尔滨工业大学(深圳); 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
QuantWM提出无需训练的2比特KV缓存量化框架,通过QSAC聚类和PSAC补偿保留注意力logits与token选择,解决视频生成中的时序闪烁,实现高达6.20倍压缩并提升视觉质量。
AI 中文摘要
KV缓存内存已成为视频生成和世界模型部署的主要瓶颈,这促使人们研究低比特量化以提高效率。现有的2比特KV缓存量化方法在VBench等视频基准上可以实现近乎无损的性能,然而,我们发现它们仍然会导致严重的时序闪烁和视觉退化。同时,更深入的研究表明,Key量化产生的重建误差小于Value,但令人惊讶的是,它却导致更大的输出退化。我们将这种差异追溯到注意力机制:小的Key扰动可以改变注意力logits,即QK^T,并改变Query所选择的时序-空间token。这些观察促使我们在KV缓存量化过程中显式地保留注意力logits和时序-空间token选择,以缓解视觉退化问题。为了解决这个问题,我们提出了QuantWM,一个无需训练且严格因果的2比特KV缓存量化框架。QuantWM引入了两种互补的技术来缓解注意力偏移。首先,量化敏感性感知聚类(QSAC)联合考虑历史Query敏感性和残差范围来选择对INT2友好的Key质心,这减少了在注意力更关键的通道中的量化误差。此外,主子空间注意力补偿(PSAC)使用低秩投影沿着主导Query子空间恢复剩余的Key误差,这提供了一种直接而有效的校正来稳定注意力logits。在Causal-Forcing、LingBot-World-v2、HY-World 1.5、Matrix-Game-2和Longcat-Video上的大量实验表明,QuantWM显著提高了视觉质量和时序一致性,同时在图像和视频质量指标上优于现有方法,实现了高达6.20倍的KV缓存内存压缩,且额外开销有限。
英文摘要
Video world models achieve long-range temporal consistency by storing KV cache during generation, but the growing cache makes KV cache memory a major deployment bottleneck, which motivates low-bit quantization study for efficiency. Existing 2-bit KV cache quantization methods can achieve nearly lossless performance on VBench, however, when applied to video world models, we find they still cause severe temporal flickering and visual degradation. Meanwhile, deeper investigates show that Key quantization produces smaller reconstruction errors than Value, but surprisingly leads to larger output degradation. We trace this discrepancy to attention in video world models: Key perturbations can change the attention logits, and shift the temporal-spatial tokens selected by Queries. These observations motivate us to preserve attention logits and temporal-spatial token selection during KV cache quantization. To address this issue, we present QuantWM, a training-free 2-bit KV cache quantization framework for video world models. QuantWM introduces two complementary techniques to mitigate the attention shifts. Firstly, quantization-sensitivity-aware clustering (QSAC) jointly considers historical Query sensitivity and residual ranges to select INT2-friendly Key centroids, which reduces quantization errors in channels that are more critical to attention. In addition, principal-subspace attention compensation (PSAC) restores the remaining Key errors along the dominant Query subspace using low-rank projections, which provides a direct and efficient correction to stabilize attention logits. Experiments on LingBot-World-v2, HY-World 1.5, Matrix-Game-2, Longcat-Video and Causal-Forcing demonstrate that QuantWM significantly improves visual quality and temporal consistency, while outperforming existing methods across benchmarks with up to 6.20 KV cache memory compression and limited additional overhead.