arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34330cs.CV

MiCo:通过语义擦除建模实现互信息覆盖优化,用于高效 MLLM 推理

MiCo: Mutual Information Coverage Optimization through Semantic Erasure Modeling for Efficient MLLM Inference

Tinghao Wang, Yichen Guo, Qizhe Zhang, Yuan Zhang, Weimin Ouyang, Rui Huang, Jiajun Cao, Sixiang Chen, Hao Jiang, Jixian Wu, Zheng Lu, Bofan Zhu, Renyuan Li, Shanghang Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

MiCo 提出一种无需训练的两阶段剪枝方法,通过语义擦除建模推导互信息覆盖目标,以贪心优化子模覆盖函数,在多种 MLLM 上实现高效推理,仅用少量视觉令牌保留高性能并显著加速。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)在多模态理解方面展现了令人印象深刻的表现,但处理大量视觉令牌会导致高昂的计算成本。尽管已有许多方法被提出以减少视觉令牌的数量,但其中大多数依赖于启发式方法,并且在剪枝过程中容易丢弃大量视觉信息,导致模型性能下降。在本工作中,通过使用语义擦除模型,我们从任务对数损失中推导出一个通用的互信息覆盖目标,并提出了 MiCo,一种无需训练的两阶段剪枝方法。MiCo 首先在视觉令牌进入语言模型之前,利用视觉信号选择一个具有代表性的候选池,然后在该候选池内进行任务感知的子集选择。在每个阶段,合适的可观测代理将推导出的目标实例化为一个单调子模覆盖函数,MiCo 在令牌预算下对此函数进行贪心优化。MiCo 在从 7B 到 13B 参数的各种 MLLM 上进行了评估,涵盖了广泛的图像和视频基准,包括通用视觉推理、细粒度 OCR 和定位、幻觉检测以及长视频理解。MiCo 在几乎所有评估模型和所有剪枝比例下都持续取得了最佳性能。在 LLaVA-NEXT-13B 上,MiCo 仅使用 5.6% 的视觉令牌,保留了 97.5% 的基线性能,并实现了 3.8 倍的推理加速。我们的实验证明了 MiCo 及其互信息覆盖目标在视觉令牌剪枝中的有效性。

英文摘要

Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visual tokens results in high computational costs. While many methods have been proposed to reduce the number of visual tokens, most of them rely on heuristics and are prone to discarding substantial visual information during pruning, leading to degradation in model performance. In this work, by using a semantic erasure model, we derive a general mutual information coverage objective from task log-loss and propose MiCo, a training-free two-stage pruning method. MiCo first uses visual signals to select a representative candidate pool before visual tokens enter the language model, then performs task-aware subset selection within it. At each stage, suitable observable proxies instantiate the derived objective as a monotone submodular coverage function, which MiCo greedily optimizes under the token budget. MiCo is evaluated on diverse MLLMs ranging from 7B to 13B parameters across a broad range of image and video benchmarks spanning general visual reasoning, fine-grained OCR and grounding, hallucination detection, and long-video understanding. MiCo consistently achieves the best performance across nearly all evaluated models under all pruning ratios. On LLaVA-NEXT-13B, MiCo uses only 5.6% visual tokens, retains 97.5% of baseline performance, and achieves a 3.8-fold inference speedup. Our experiments demonstrate the effectiveness of MiCo and our mutual information coverage objective for visual token pruning.

发表机构

  • Peking University(北京大学)
  • University of Electronic Science and Technology of China(电子科技大学)
  • Nanyang Technological University(南洋理工大学)
  • The University of Hong Kong(香港大学)
  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑