AKVQ-VL:面向视觉语言模型的注意力感知KV缓存自适应2比特量化
AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models
- Tsinghua University(清华大学)
- Huawei Technologies Co., Ltd(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出AKVQ-VL,一种面向视觉语言模型的自适应2比特KV缓存量化方法,利用文本和枢轴token显著注意力模式分配比特预算并结合WHT处理异常值,在12个任务上保持或提升准确率,显著降低内存并提升吞吐量。
AI中文摘要:
视觉语言模型(VLM)在多模态任务中表现出色。然而,过长的多模态输入会导致过大的键值(KV)缓存,从而引发显著的内存消耗和I/O瓶颈。针对大型语言模型(LLM)的先前KV量化方法可能缓解这些问题,但忽略了多模态token的注意力显著性差异,导致性能次优。本文研究了VLM中的注意力感知token显著性模式,并提出了AKVQ-VL。AKVQ-VL利用提出的文本显著注意力(TSA)和枢轴token显著注意力(PSA)模式自适应分配比特预算。此外,实现极低比特量化需要有效处理KV张量中的异常值。AKVQ-VL利用Walsh-Hadamard变换(WHT)构建无异常值的KV缓存,从而降低量化难度。在12个长上下文和多模态任务上的2比特量化评估表明,AKVQ-VL保持甚至提高了准确率,优于面向LLM的方法。AKVQ-VL可将峰值内存使用量降低2.13倍,支持多达3.25倍的更大批处理大小和2.46倍的吞吐量。
英文摘要:
Vision-language models (VLMs) show remarkable performance in multimodal tasks. However, excessively long multimodal inputs lead to oversized Key-Value (KV) caches, resulting in significant memory consumption and I/O bottlenecks. Previous KV quantization methods for Large Language Models (LLMs) may alleviate these issues but overlook the attention saliency differences of multimodal tokens, resulting in suboptimal performance. In this paper, we investigate the attention-aware token saliency patterns in VLM and propose AKVQ-VL. AKVQ-VL leverages the proposed Text-Salient Attention (TSA) and Pivot-Token-Salient Attention (PSA) patterns to adaptively allocate bit budgets. Moreover, achieving extremely low-bit quantization requires effectively addressing outliers in KV tensors. AKVQ-VL utilizes the Walsh-Hadamard transform (WHT) to construct outlier-free KV caches, thereby reducing quantization difficulty. Evaluations of 2-bit quantization on 12 long-context and multimodal tasks demonstrate that AKVQ-VL maintains or even improves accuracy, outperforming LLM-oriented methods. AKVQ-VL can reduce peak memory usage by 2.13x, support up to 3.25x larger batch sizes and 2.46x throughput.