像素可解码性并非压缩信号:对视觉KV缓存驱逐的重要性代理进行因果评估
Pixel Decodability Is Not a Compression Signal: Causally Evaluating Importance Proxies for Visual KV-Cache Eviction
浏览论文内容
中文总结 AI 辅助
本文通过因果评估证明视觉KV缓存中像素可解码保留是任务惰性的,不能作为有效的压缩信号,而注意力是唯一微弱但显著的重要性代理。
中文摘要 AI 辅助
视觉语言模型在其视觉键值缓存中保留了大量的像素可解码视觉内容。我们表明,在我们的设定中,这种保留是任务惰性的:在我们预先注册的测试中,一个单元保留多少从不正向追踪回答问题所需的计算是否因果依赖于它。我们使用学习到的像素反演解码器测量保留量,并使用单超补丁KV消融(即教师强制的黄金答案对数概率下降)测量因果使用,并在预先注册、符号校准、留出设计下,在图像内关联两者。保留量与注意力解耦,并且在功效充足的零假设下,与因果利用解耦。利用并非对所有代理都惰性:注意力微弱但显著地追踪它,这是我们发现的唯一能做到这一点的信号,也是设计的阳性对照。我们将像素可解码保留描述为视觉KV缓存的一个信息轴,与功能轴正交。在我们的模型对中,缓存包含的任务惰性内容量因架构而异:无编码器模型保留的量是基于编码器模型的2.7倍。工程后果是一个受控的阴性结果。在超补丁粒度下,去混杂的像素可解码保留对KV驱逐的排序不优于随机;在令牌粒度下,在较大预算下仅获得微弱的逆重要性信号,且在每种预算下都被注意力幅度主导。在我们的设定中,像素可解码可重构性在我们测试的任何粒度下都不是有竞争力的KV压缩信号。
英文摘要
Vision-language models retain a substantial amount of pixel-decodable visual content in their visual key-value cache. We show, in our setting, that this retention is task-inert: across our preregistered tests, how much a unit retains never positively tracks whether the computation that answers the question causally relies on it. We measure retention with a learned pixel-inversion decoder and causal use with single-super-patch KV ablation, the teacher-forced drop in gold-answer log-probability, and relate the two within images under a preregistered, sign-calibrated, held-out design. Retention is decoupled from attention and, in a well-powered null, from causal utilization. Utilization is not inert to every proxy: attention weakly but significantly tracks it, the only signal we find that does and the design's positive control. We characterize pixel-decodable retention as an informational axis of the visual KV cache, orthogonal to the functional one. How much task-inert content a cache holds differs by architecture in our model pair: the encoder-free model retains 2.7 times more than the encoder-based one. The engineering consequence is a controlled negative result. At super-patch granularity, deconfounded pixel-decodable retention ranks KV eviction no better than random; at token granularity it acquires only a weak inverse-importance signal at larger budgets, dominated at every budget by attention magnitude. In our setting, pixel-decodable reconstructability is not a competitive KV-compression signal at any granularity we test.
发表机构
- Institute of Science Tokyo(东京科学大学)
- Zhejiang University(浙江大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。