发表机构
CoreAI, Microsoft(微软CoreAI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HeadGuard通过固定高精度掩码保护约1/8的KV头部,在八个VLM和八个基准上显著恢复低比特量化损失,将2比特判别性平均分从0.436提升至0.580。
AI 中文摘要
低比特键值(KV)缓存量化节省存储,但会显著降低视觉语言模型(VLM)的准确性。我们提出HeadGuard,一种可组合的头部保护方法,通过固定的高精度掩码增强基础KV缓存量化器。图像敏感度和输出敏感度分数在离线状态下选择物理KV头部,主要实验中约1/8的头部受到保护;其图像键和可选值保持为bfloat16(BF16)格式,而基础量化器对未受保护的图像条目进行量化。在八个VLM、三个基础量化器和八个基准(六个判别性任务和两个生成性任务)上,HeadGuard在较弱的量化器上恢复了大部分丢失的准确性,其中对Qwen和InternVL的提升最为显著。在2比特设置下,八个模型在六个判别性任务上的平均分数从最弱基础量化器的0.436提升至0.580;保护还能改善生成答案和字幕与BF16输出的一致性。在两种测试的校准数据集上,所有三个量化器的平均准确性提升均持续存在。仅保护键在较低建模存储成本下仍能保持显著的恢复效果。通过模拟量化评估,HeadGuard提供了一种可组合的方式,在不替换底层量化器的情况下提高低比特VLM的准确性。
英文摘要
Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composable head-protection method that augments a base KV-cache quantizer with a fixed high-precision mask. Image-sensitivity and output-sensitivity scores select physical KV heads offline, with approximately 1/8 protected in the main experiments; their image keys and optionally values remain in bfloat16 (BF16), while the base quantizes unprotected image entries. Across eight VLMs, three base quantizers, and eight benchmarks (six discriminative and two generative), HeadGuard recovers a substantial fraction of lost accuracy on weaker quantizers, with the strongest gains for Qwen and InternVL. At 2 bits, the six-task discriminative mean over eight models rises from 0.436 to 0.580 on the weakest base; protection can also improve generated answers and caption fidelity to BF16 outputs. Mean accuracy gains persist across all three quantizers with both tested calibration datasets. Keys-only protection retains substantial recovery at lower modeled storage cost. Evaluated through simulated quantization, HeadGuard offers a composable way to improve low-bit VLM accuracy without replacing the underlying quantizer.