ECOKV:基于互补多样性度量的几何感知KV缓存驱逐
ECOKV: Geometry-Aware KV Cache Eviction via Complementary Diversity Metrics
浏览论文内容
中文总结 AI 辅助
针对多模态大模型KV缓存内存开销问题,提出ECOKV方法,结合欧氏距离与余弦相似度的几何感知复合度量及自适应加权,缩减观察窗口,实现高效缓存驱逐并达到最先进性能。
中文摘要 AI 辅助
尽管多模态大语言模型(MLLMs)在多种任务中表现出色,但其可扩展性仍受限于KV缓存存储的内存和计算开销。现有的KV缓存驱逐方法将基于余弦相似度的多样性度量与重要性度量相结合,以选择性地保留关键的键值对。然而,余弦相似度涉及归一化,会丢弃幅度信息,并且由于隐藏表示的各向异性特性,它往往在各层产生均匀的高相似度值。在我们的ECOKV研究中,我们严格解构了现有多样性度量的能力。超越简单的测量,我们提出了一种几何感知的复合度量,联合利用欧几里得距离和余弦相似度,从互补的角度捕捉令牌多样性。此外,我们使用这两种度量来估计每个注意力头的冗余水平,从而在令牌选择过程中实现多样性与重要性分数之间的自适应加权。最后,我们证明通常用于保留最近令牌的观察窗口可以大幅缩减,从而将更多缓存容量分配给信息丰富的令牌,并带来一致的改进。大量实验表明,ECOKV在各种压缩比率下均达到最先进性能,并可无缝集成到现有的KV缓存驱逐方法中。我们进一步分析了重要性与多样性之间的关系,并考察了跨层和注意力头的冗余模式。
英文摘要
Although multimodal Large Language Models (MLLMs) excel in diverse tasks, their scalability remains limited by the memory and computational overhead of KV cache storage. Recent KV cache eviction approaches incorporate a cosine similarity-based diversity metric with importance metrics to selectively retain critical key-value pairs. However, cosine similarity involves normalization that discards magnitude information, and it often yields uniformly high similarity values across layers due to the anisotropy property of hidden representations. In our study ECOKV, we rigorously deconstruct the capabilities of existing diversity metrics. Moving beyond simple measurement, we propose a geometry-aware composite metric that jointly leverages Euclidean distance and cosine similarity to capture token diversity from complementary perspectives. Furthermore, we use these two metrics to estimate the redundancy level of each attention head, allowing adaptive weighting between diversity and importance scores during token selection. Finally, we demonstrate that the observation window commonly employed to preserve recent tokens can be substantially reduced, thereby allocating more cache capacity to informative tokens and yielding consistent improvements. Extensive experiments demonstrate that ECOKV achieves state-of-the-art performance under various compression ratios and can be seamlessly integrated with existing KV cache eviction methods. We further analyze the relationship between importance and diversity, and examine redundancy patterns across layers and attention heads.
发表机构
- MediaTek Inc.(联发科股份有限公司)
- National Taiwan University(台湾大学)
机构由 AI 辅助整理,请以论文原文为准。