CONDUIT:一种用于视觉语言模型中KV缓存复用的统一残差流恢复框架
CONDUIT: A Unified Residual-Stream Restoration Framework for KV Cache Reuse in Vision-Language Models
浏览论文内容
中文总结 AI 辅助
CONDUIT提出一种无需训练的KV缓存刷新策略,通过残差流恢复统一单/多图像重用,在10%预算下达到全预填充97%以上性能并显著加速。
中文摘要 AI 辅助
视觉语言模型(VLM)经常回答关于重复出现的视觉内容的新问题,此时重用键值(KV)缓存可以避免重新编码昂贵的视觉前缀。然而,当相同的视觉内容在改变的前缀下出现时,精确前缀重用会失败。选择性重计算可以在小的视觉令牌预算下恢复质量,但仅当刷新正确的陈旧令牌时。原始注意力选择可能会将预算浪费在高注意力但值范数代理分数低的令牌上,以及浪费在与查询无关的图像上。为解决这些失败模式,我们提出CONDUIT,一种无需训练的刷新策略,将单图像和多图像重用统一为残差流恢复。基于范数加权注意力,CONDUIT使用缓存键查询注意力和可访问的预输出缓存值范数代理对缓存的视觉令牌进行排名,然后在一次全局选择之前应用经验性的图像级相关性放大。对于单图像,系数为1,规则简化为图像内令牌选择。该方法保留模型架构和权重,仅在推理时增加一次查询条件评分传递。在10%的刷新预算下,CONDUIT在三个VLM骨干网络上实现了对应全预填充五数据集平均值的97.0-99.5%,并在平均上领先于预算方法;在MMLongBench-Doc延迟子集上,它使用全预填充FLOPs的13.5%,并实现了2.99倍的首次令牌时间加速。
英文摘要
Vision-language models (VLMs) often answer new questions about recurring visual content, where reusing the key-value (KV) cache can avoid re-encoding expensive visual prefixes. Exact-prefix reuse, however, fails when the same visual content appears under a changed prefix. Selective recomputation can recover quality under a small visual-token budget, but only when the right stale tokens are refreshed. Raw-attention selection can waste budget on high-attention tokens with small value-norm proxy scores and on query-irrelevant images. To address these failure modes, we propose CONDUIT, a training-free refresh policy that unifies single- and multi-image reuse as residual-stream restoration. Building on norm-weighted attention, CONDUIT ranks cached visual tokens using cached-key query attention and an accessible pre-output cached-value-norm proxy, then applies empirical image-level relevance amplification before one global selection. With one image, the coefficient is one and the rule reduces to intra-image token selection. The method preserves model architecture and weights, adding only a single query-conditioned scoring pass at inference. At a 10% refresh budget, CONDUIT achieves 97.0-99.5% of the corresponding full-prefill five-dataset average across three VLM backbones and leads budgeted methods on average; on the MMLongBench-Doc latency subset, it uses 13.5% of full-prefill FLOPs and achieves a 2.99x time-to-first-token speedup.
发表机构
- The Chinese University of Hong Kong(香港中文大学)
- Fudan University(复旦大学)
- LIGHTSPEED
机构由 AI 辅助整理,请以论文原文为准。