学习在多模态大语言模型(MLLMs)中预测中间层注意力以进行视觉token剪枝
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Pruning
浏览论文内容
中文总结 AI 辅助
该研究针对MLLMs视觉token剪枝的固定层非最优及计算成本高的问题,提出MAP方法,实现仅保留5.56%视觉token时维持97.5%性能,获3.09倍端到端加速。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)在各类视觉-语言任务中表现出色,但处理大量视觉token的成本限制了其效率。视觉token剪枝可降低该成本,但需要准确的token重要性估计。近期研究表明,来自中间语言模型层的文本到视觉注意力能有效指导视觉token剪枝,通常使用预定义中间层的注意力来选择要保留的视觉token。因此仍存在两个问题:其一,我们的分析显示,对问题最敏感的注意力层会随样本显著变化,固定层并非最优;其二,获取合适中间层的注意力需要通过多个语言模型层处理大量视觉token,此时已消耗大量计算资源。为解决这两个问题,我们提出中间层注意力预测(MAP)方法,该方法利用问题对比教师选择,通过对比原始问题与参考问题下的注意力,识别样本特定的教师层,并将所选层的注意力蒸馏为轻量预测器,用于从多模态输入特征估计视觉token重要性。推理阶段,MAP将预测的重要性分数与多样性准则结合,在第一个语言模型层之前剪枝视觉token。因此,MAP无需注意力图即可进行剪枝,且与现有推理加速技术兼容。在LLaVA-NeXT-7B的10个基准测试中,MAP仅保留5.56%的视觉token,却维持了未剪枝模型97.5%的性能,实现了3.09倍的端到端加速。
英文摘要
Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual tokens. Visual token pruning can reduce this cost, but requires accurate token importance estimates. Recent studies have demonstrated that text-to-vision attention from middle language model layers can effectively guide visual token pruning, typically using attention from a predefined middle layer to select the visual tokens to retain. Two problems therefore remain. First, our analysis shows that the layer whose attention is most responsive to the question varies substantially across samples, making a fixed layer suboptimal. Second, obtaining attention from the appropriate middle layer requires processing numerous visual tokens through several language model layers, by which point considerable computation has already been spent. To address both problems, we propose Middle-layer Attention Prediction (MAP), which uses Question Contrastive Teacher Selection to identify a sample-specific teacher layer by contrasting attention under the original and reference questions, and distills attention from the selected layer into a lightweight predictor that estimates visual token importance from multi-modal input features. During inference, MAP combines the predicted importance scores with a diversity criterion to prune visual tokens before the first language model layer. Thus, MAP requires no attention maps for pruning and remains compatible with existing inference acceleration techniques. Across ten benchmarks on LLaVA-NeXT-7B, MAP retains 97.5% of the unpruned model performance with only 5.56% of the visual tokens, yielding a 3.09x end-to-end speedup.