LFTR:面向多模态大语言模型的无学习令牌缩减方法
LFTR: Learning-Free Token Reduction for Multimodal Large Language Models
AI总结:
本文提出无需训练的LFTR视觉令牌缩减方法,可适配多种开源MLLM,最高缩减16倍视觉令牌且保持或提升视觉问答性能,并可与其他加速技术互补。
AI中文摘要:
多模态大语言模型(MLLM)已在各类多模态任务中取得卓越成功,但其部署往往受到巨大计算需求和漫长推理时间的限制。鉴于视觉模态通常包含比文本模态更全面的信息,编码后的表示会包含大量令牌,而注意力机制的二次复杂度会带来显著的计算开销。现有令牌缩减方法通常局限于特定模型架构,且往往需要大量重新训练或微调,限制了其在许多最先进模型上的适用性。本文提出一种面向MLLM的无学习令牌缩减(LFTR)方法。LFTR可无缝集成到大多数开源MLLM架构中,无需额外微调。该方法利用视觉表示中的冗余,在有效缩减令牌的同时保持MLLM的通用推理性能。我们在多种MLLM架构(LLaVA、MiniGPT、QwenVL)上进行实验,结果表明LFTR可将视觉令牌最多缩减16倍,并在无学习设置下于主流视觉问答基准上保持甚至提升性能。此外,LFTR与视觉编码器压缩、训练后量化等其他加速技术互补,可进一步推动MLLM的高效部署。项目位于https://anonymous.4open.science/r/LFTR-AAAI-0528。
英文摘要:
Multimodal Large Language Models (MLLMs) have demonstrated exceptional success in various multimodal tasks, yet their deployment is frequently limited by substantial computational demands and prolonged inference times. Given that the vision modality typically contains more comprehensive information than the text modality, resulting in encoded representations comprising an extensive number of tokens, leading to significant computational overhead due to the quadratic complexity of the attention mechanism. Current token reduction methods are typically restricted to specific model architectures and often necessitate extensive retraining or fine-tuning, restricting their applicability to many state-of-the-art models. In this paper, we introduce a learning-free token reduction (LFTR) method designed for MLLMs. LFTR can be seamlessly integrated into most open-source MLLM architectures without requiring additional fine-tuning. By capitalizing on the redundancy in visual representations, our approach effectively reduces tokens while preserving the general inference performance of MLLMs. We conduct experiments on multiple MLLM architectures (LLaVA, MiniGPT, QwenVL), and our results show that LFTR achieves up to a $16\times$ reduction of visual tokens while maintaining or even enhancing performance on mainstream vision question-answering benchmarks, all in a learning-free setting. Additionally, LFTR is complementary to other acceleration techniques, such as vision encoder compression and post-training quantization, further promoting the efficient deployment of MLLMs. Our project is available at https://anonymous.4open.science/r/LFTR-AAAI-0528.