arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16326cs.CV

CRISP:用于高效LVLM推理的预LLM文本驱动视觉令牌剪枝

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference

  • College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)
  • School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院)
  • School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院)
  • College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)

机构由 AI 辅助整理,请以论文原文为准。

Xu Li, Yi Zheng, Mengyang Zhao, Yuxuan Liang, Zhe Liu, Rui Zhu, Xiaolei Chen, Wei Zhou, Baoquan Zhao, Juncen Guo

AI总结:

针对大型视觉语言模型推理开销大的问题,提出CRISP框架,通过文本驱动在预LLM阶段剪枝视觉令牌,分两阶段工作,实验表明其在激进剪枝率下能保持高性能,降低推理成本和延迟,是高效LVLM推理的实用方案。

AI中文摘要:

大型视觉语言模型(LVLMs)通常需要处理数百到数千个视觉令牌,导致大量推理开销。现有视觉令牌剪枝方法要么在LLM之前使用与文本无关的启发式方法,要么在LLM内部以效率和嘈杂的跨模态注意力为代价进行剪枝。为解决这些限制,我们提出CRISP,一种预LLM但文本驱动的视觉令牌剪枝框架,它保留与指令相关的证据和基本场景上下文。CRISP在两阶段管道中工作:第一阶段首先识别文本对齐的视觉令牌,第二阶段通过语义多样性增强上下文完整性。在LLaVA-1.5和LLaVA-NeXT上的大量实验表明,CRISP在激进剪枝率下实现了卓越的性能保留,在将推理成本和延迟降低两倍以上的同时,保持高达99.5%的准确率。CRISP是高效LVLM推理的实用解决方案,尤其是在资源受限场景中。

英文摘要:

Large Vision-Language Models (LVLMs) typically require processing hundreds to thousands of visual tokens, leading to substantial inference overhead. Existing visual token pruning methods either operate before the LLM using text-agnostic heuristics or prune inside the LLM at the cost of efficiency and noisy cross-modal attention. To address these limitations, we propose CRISP, a pre-LLM yet text-driven visual token pruning framework that preserves both instruction-relevant evidence and essential scene context. CRISP works in a two-stage pipeline: Stage 1 first identifies text-aligned visual tokens, and Stage 2 enhances contextual completeness through semantic diversity. Extensive experiments on LLaVA-1.5 and LLaVA-NeXT demonstrate that CRISP achieves superior performance retention under aggressive pruning ratios, maintaining up to 99.5% accuracy while reducing inference cost and latency by more than 2 times. CRISP serves as a practical solution for efficient LVLM inference, especially in resource-constrained scenarios.

补充信息

↑