arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MWOP:面向高效多模态大语言模型的模态感知宽度方向操作剪枝

MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs

Xudong Wang, Hao Wu, Haozhe Hu, Peiran Yin, Xinghao Chen, Yunpu Ma, Wei Zhang, Xiaoyu Shen

arXiv 2610.01434首次发表:更新:

发表机构

Eastern Institute of Technology; Shanghai Jiao Tong University; The Hong Kong Polytechnic University; LMU Munich(东方理工高等研究院; 上海交通大学; 香港理工大学; 慕尼黑大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态大语言模型推理成本高的问题,提出模态感知宽度方向操作剪枝(MWOP),通过细粒度剪枝注意力路径和FFN通道,在保持性能的同时实现1.6倍加速,并与令牌压缩方法互补。

AI 中文摘要

多模态大语言模型(MLLMs)在处理长视觉-文本序列时会产生大量的推理成本。现有的操作压缩方法虽然利用了模态级别的冗余,但它们在很大程度上将注意力头内和共享前馈网络(FFN)通道内的计算视为统一单元,留下了更细粒度的冗余未被充分探索。我们发现,冗余在同一注意力头内的不同模态交互路径之间以及同一FFN通道的视觉和文本执行之间均存在差异。基于这些发现,我们提出了模态感知宽度方向操作剪枝(MWOP),该方法独立地剪枝每一层内的视觉到视觉(V2V)、文本到视觉(T2V)和文本到文本(T2T)注意力路径,并分别针对视觉和文本输入选择FFN通道。一阶泰勒准则指导剪枝过程,在注意力剪枝后重新评估FFN重要性,并采用基于LoRA的恢复训练。为了将由此产生的细粒度稀疏性转化为实际加速,我们进一步开发了路径稀疏的Triton注意力内核和紧凑的视觉侧FFN执行。MWOP在保留令牌序列的同时减少了注意力和FFN计算,使其与令牌压缩互补,并能够同时减少序列长度和每个令牌的计算量。在LLaVA-OneVision-7B上,仅MWOP就实现了1.6倍的预填充加速,在12个基准测试中平均性能保持率为99.7%。与两种代表性的令牌压缩方法结合后,它们的预填充加速分别从2.0倍和1.9倍提高到2.9倍和2.7倍。在Qwen2.5-VL-7B上的结果进一步证明了其跨架构的适用性。代码可在以下网址获取:此https URL。

英文摘要

Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual executions of the same FFN channel. Based on these findings, we propose Modality-aware Width-wise Operation Pruning (MWOP), which independently prunes visual-to-visual (V2V), text-to-visual (T2V), and text-to-text (T2T) attention paths within each layer, and separately selects FFN channels for visual and textual inputs. A first-order Taylor criterion guides the pruning process, with FFN importance re-evaluated after attention pruning and LoRA-based recovery training. To translate the resulting fine-grained sparsity into practical acceleration, we further develop path-sparse Triton attention kernels and compact visual-side FFN execution. MWOP preserves the token sequence while reducing attention and FFN computation, making it complementary to token compression and enabling simultaneous reduction of sequence length and per-token computation. On LLaVA-OneVision-7B, MWOP alone achieves a $1.6\times$ prefill speedup with 99.7\% average performance retention across 12 benchmarks. Combined with two representative token compression methods, it further increases their prefill speedups from $2.0\times$ and $1.9\times$ to $2.9\times$ and $2.7\times$, respectively. Results on Qwen2.5-VL-7B further demonstrate its applicability across architectures. The code is available at https://github.com/EIT-NLP/MWOP.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑