发表机构
Northwest Polytechnical University; Intellifusion(西北工业大学; 云天励飞)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
P4Q提出视觉令牌剪枝与低比特量化的协同设计框架,通过量化感知令牌选择和剪枝感知校准,在LLaVA-NeXT上实现2.8倍端到端加速并保持更高精度。
AI 中文摘要
视觉语言模型在广泛的多模态应用中取得了强劲的性能,但其巨大的计算和内存成本阻碍了高效部署。视觉令牌剪枝和训练后量化分别从序列长度和数值精度这两个互补维度降低推理开销。现有工作流程通常独立优化这些技术或顺序应用它们。它们不同的优化目标使得关键交互未被处理,并限制了可实现的压缩性能。我们重新审视这些设计,提出了P4Q,一个实用的协同设计框架,联合优化视觉令牌剪枝和低比特量化以实现高效的视觉语言模型推理。首先,P4Q在大语言模型之前引入了一种量化感知的视觉令牌选择策略。它对投影器产生的特征副本应用伪量化,并使用从这些伪量化特征计算得到的统计量来选择视觉令牌,从而将选择器的基于特征的决策条件化为模拟的低比特扰动。其次,P4Q引入了一种剪枝感知的量化校准策略。它使用与剪枝相同的选择策略,在保留令牌分布上校准量化模型,从而使校准过程与部署时使用的剪枝执行路径对齐。通过耦合这两个组件,P4Q在保持可比任务性能的同时实现了显著的推理加速,从而比独立优化的流程获得了更好的效率-精度权衡。例如,在LLaVA-NeXT上,P4Q在八个不同的测试集上实现了平均2.8倍的端到端推理加速,同时保持了比先前压缩和量化方法更高的精度。
英文摘要
Vision language models have achieved strong performance across a wide range of multimodal applications, yet their substantial computational and memory costs hinder efficient deployment. Visual token pruning and post-training quantization reduce inference overhead along two complementary dimensions, namely sequence length and numerical precision. Existing workflows typically optimize these techniques independently or apply them sequentially. Their distinct optimization objectives leave critical interactions unaddressed and constrain the achievable compression performance. We revisit these designs and present P4Q, a practical co-design framework that jointly optimizes visual token pruning and low-bit quantization for efficient VLM inference. First, P4Q introduces a quantization-aware visual token selection strategy before the LLM. It applies fake quantization to copies of the features produced by the projector and selects visual tokens using statistics computed from these fake-quantized features, thereby conditioning the selector's feature-based decisions on simulated low-bit perturbations. Second, P4Q introduces a pruning-aware quantization calibration strategy. It uses the same selection strategy as pruning to calibrate the quantized model on the retained-token distribution, thereby aligning the calibration process with the pruned execution path used during deployment. By coupling these two components, P4Q achieves substantial inference speedups while maintaining comparable task performance, resulting in a better efficiency-accuracy trade-off than independently optimized pipelines. For instance, on LLaVA-NeXT, P4Q achieves an average end-to-end inference speedup of 2.8x across eight distinct test sets, while retaining higher accuracy than prior compression and quantization methods.