发表机构
Shenzhen International Center for Industrial and Applied Mathematics; Shenzhen Research Institute of Big Data; The Chinese University of Hong Kong, Shenzhen; Shenzhen Loop Area Institute(深圳国际工业与应用数学中心; 深圳大数据研究院; 香港中文大学(深圳); 深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VisCache是一种无需训练的即插即用视觉KV缓存剪枝框架,通过两阶段协同操作提升VLLM长上下文推理效率,实现最高2.35倍加速,建立效率与性能的新帕累托前沿。
AI 中文摘要
视觉大语言模型(VLLMs)在多模态推理中取得了显著成功,但其长上下文推理因视觉键值(KV)缓存的巨大计算与内存开销而成本过高。现有KV压缩方法常对视觉 token 和层采用统一剪枝,导致大量信息丢失与性能下降。为解决该挑战,我们提出VisCache,即一种无需训练、即插即用的粗到细视觉KV缓存剪枝框架,由两个协同阶段构成:第一,轻量VLM通过选择性转发语义信息丰富的关键帧过滤时间冗余;第二,我们引入PruneKV,一种专为VLLMs注意力动态设计的外科手术式KV压缩算法。与刚性剪枝策略不同,PruneKV采用抛物线式层预算分配及非对称更新机制,在剪枝键的同时融合值,从而保留关键上下文信息。大量实验表明,VisCache大幅提升推理效率,实现最高2.35倍的加速与显著的内存减少,同时在仅保留19%-28% KV缓存的情况下保持竞争力性能,且始终优于现有基线,为长上下文VLLM推理建立了效率与性能间的新帕累托前沿。代码可在提供的链接获取。
英文摘要
While Vision Large Language Models (VLLMs) have achieved remarkable success in multimodal reasoning, their long-context inference remains prohibitively expensive due to the massive computation and memory overhead of visual Key-Value (KV) caches. Existing KV compression methods often apply uniform pruning across visual tokens and layers, leading to substantial information loss and degraded performance.To address this challenge, we propose \textbf{VisCache}, a plug-and-play framework for coarse-to-fine \textbf{Vis}ual KV \textbf{Cache} pruning without training, which consists of two synergistic stages. First, a lightweight VLM filters temporal redundancy by selectively forwarding semantically informative keyframes. Second, we introduce {PruneKV}, a surgical KV compression algorithm tailored to the attention dynamics of VLLMs. Unlike rigid pruning strategies, PruneKV adopts a parabolic layer-wise budget allocation together with an asymmetric update mechanism that selectively prunes keys while fusing values, thereby preserving critical contextual information. Extensive experiments demonstrate that VisCache substantially improves inference efficiency, achieving up to {2.35$\times$ speedup} and significant memory reduction while maintaining competitive performance with only {19--28\%} KV cache retention. VisCache consistently outperforms existing baselines, establishing a new Pareto frontier between efficiency and performance for long-context VLLM inference. Code is available at https://github.com/Wlklk/VisCache