CRISP:用于高效LVLM推理的预LLM文本驱动视觉令牌剪枝
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference
- College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)
- School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院)
- School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院)
- College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大型视觉语言模型推理开销大的问题,提出CRISP框架,通过文本驱动在预LLM阶段剪枝视觉令牌,分两阶段工作,实验表明其在激进剪枝率下能保持高性能,降低推理成本和延迟,是高效LVLM推理的实用方案。
AI中文摘要:
大型视觉语言模型(LVLMs)通常需要处理数百到数千个视觉令牌,导致大量推理开销。现有视觉令牌剪枝方法要么在LLM之前使用与文本无关的启发式方法,要么在LLM内部以效率和嘈杂的跨模态注意力为代价进行剪枝。为解决这些限制,我们提出CRISP,一种预LLM但文本驱动的视觉令牌剪枝框架,它保留与指令相关的证据和基本场景上下文。CRISP在两阶段管道中工作:第一阶段首先识别文本对齐的视觉令牌,第二阶段通过语义多样性增强上下文完整性。在LLaVA-1.5和LLaVA-NeXT上的大量实验表明,CRISP在激进剪枝率下实现了卓越的性能保留,在将推理成本和延迟降低两倍以上的同时,保持高达99.5%的准确率。CRISP是高效LVLM推理的实用解决方案,尤其是在资源受限场景中。
英文摘要:
Large Vision-Language Models (LVLMs) typically require processing hundreds to thousands of visual tokens, leading to substantial inference overhead. Existing visual token pruning methods either operate before the LLM using text-agnostic heuristics or prune inside the LLM at the cost of efficiency and noisy cross-modal attention. To address these limitations, we propose CRISP, a pre-LLM yet text-driven visual token pruning framework that preserves both instruction-relevant evidence and essential scene context. CRISP works in a two-stage pipeline: Stage 1 first identifies text-aligned visual tokens, and Stage 2 enhances contextual completeness through semantic diversity. Extensive experiments on LLaVA-1.5 and LLaVA-NeXT demonstrate that CRISP achieves superior performance retention under aggressive pruning ratios, maintaining up to 99.5% accuracy while reducing inference cost and latency by more than 2 times. CRISP serves as a practical solution for efficient LVLM inference, especially in resource-constrained scenarios.