发表机构
Zhejiang University; Ant Group; The State Key Laboratory of Blockchain and Data Security, Zhejiang University; Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(浙江大学; 蚂蚁集团; 浙江大学区块链与数据安全国家重点实验室; 杭州高新技术产业开发区(滨江)区块链与数据安全研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多模态大语言模型中KV缓存问题,提出MM-ShiftKV方法,通过构建方差扩展查询代理近似解码时查询行为,基于聚合注意力质量估计KV重要性,在严格缓存预算下性能优于现有方法。
AI 中文摘要
键值(KV)缓存对于多模态大语言模型(MLLMs)的高效推理至关重要,但其内存占用随上下文长度线性增长,因大量视觉令牌成为主要瓶颈。近期预填充阶段KV选择方法从预填充统计估计KV重要性,隐含假设预填充时查询代表解码时情况。我们表明此假设在多模态推理中不成立,解码时查询方差比预填充阶段表示大得多,导致在缓存预算紧张时KV重要性估计不稳定。小排序错误会过度丢弃语义关键视觉令牌并降低基础和推理性能。我们提出MM-ShiftKV,一种无需训练、解码感知且仅严格用于预填充的KV选择方法。MM-ShiftKV通过构建方差扩展查询代理在预填充期间近似解码时查询行为,并基于聚合注意力质量估计提示KV重要性。多模态基准实验表明,在严格KV缓存预算下,MM-ShiftKV始终优于现有方法。
英文摘要
Key-Value (KV) caching is essential for efficient inference in multimodal large language models (MLLMs), yet its memory footprint grows linearly with context length and becomes a major bottleneck due to the large number of visual tokens. Recent prefill-stage KV selection methods estimate KV importance from prefilling statistics, implicitly assuming that prefilling-time queries are representative of those encountered during decoding. We show that this assumption breaks down in multimodal inference, where decoding-time queries exhibit substantially larger variance than prefilling-stage representations, leading to unstable KV importance estimation under tight cache budgets. As a result, small ranking errors can disproportionately discard semantically critical visual tokens and degrade grounding and reasoning performance. We propose MM-ShiftKV, a training-free, decode-aware and strictly prefill-only KV selection method. MM-ShiftKV approximates decoding-time query behavior during prefilling by constructing variance-expanded query proxies and estimates prompt KV importance based on their aggregated attention mass. Experiments on multimodal benchmarks demonstrate that MM-ShiftKV consistently outperforms existing methods under strict KV-cache budgets. Our code is available at https://github.com/zjuDBxAI/MM-ShiftKV.
Comments19 pages, 11 figures