发表机构
Xiamen University; Xiamen Ocean Vocational College; Sino-Russian Research Center for Digital Economy(厦门大学; 厦门海洋职业技术学院; 中俄数字经济研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多模态大语言模型现有无训练剪枝方法忽略令牌重要性动态变化的问题,提出趋势感知剪枝框架,通过捕捉注意力流动量实现动态修正,大幅减少视觉令牌且保持性能,实现高效多模态推理。
AI 中文摘要
视觉令牌剪枝对于高效多模态大语言模型(Multimodal Large Language Models, MLLMs)至关重要,但现有无需训练的方法存在关键局限:它们依赖静态瞬时启发式规则执行不可逆过滤,忽略了MLLMs的层级特性——令牌重要性常随层动态变化而非固定不变,导致深层推理必需的令牌常被浅层估计过早丢弃。为解决该问题,我们提出Trend-aware Pruning(趋势感知剪枝)框架,将剪枝从局部快照决策升级为时间轨迹建模问题。该方法不依赖孤立分数,而是捕捉注意力流的动量,实现动态修正机制,选择性重新激活“后发”令牌——那些初始被低估但语义重要性上升的令牌,从而避免关键视觉线索丢失。大量实验表明,该方法在各类多模态任务中实现了更优的效率-性能权衡,尤其可减少77.8%以上的视觉令牌,最终层仅保留约23个令牌,同时维持竞争力性能,为高效率多模态推理提供了稳健且可逆的解决方案。
英文摘要
While visual token pruning is essential for efficient Multimodal Large Language Models (MLLMs), existing training-free methods suffer from a critical limitation: they rely on static, instantaneous heuristics to perform irreversible filtering. This approach ignores the hierarchical nature of MLLMs, where token importance often evolves dynamically rather than remaining fixed across layers. Consequently, tokens essential for deep-layer reasoning are often prematurely discarded by shallow-layer estimates. To address this, we propose Trend-aware Pruning, a novel framework that elevates pruning from a local snapshot decision to a temporal trajectory modeling problem. Instead of relying on isolated scores, our method captures the momentum of attention flow. This enables a dynamic rectification mechanism that selectively reactivates "late-blooming" tokens, those initially undervalued but exhibiting rising semantic importance, thereby preventing the loss of critical visual cues. Extensive experiments demonstrate that our approach achieves a superior efficiency-performance trade-off across diverse multimodal tasks. Notably, it reduces visual tokens by over 77.8%, retaining only approximately 23 tokens in the final layer while maintaining competitive performance, offering a robust and reversible solution for high-efficiency multimodal inference.