GSTEP:面向高效视频大语言模型的全局时空密度驱动视觉令牌剪枝
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models
浏览论文内容
中文总结 AI 辅助
本文提出即插即用的GSTEP剪枝框架,解决现有视频令牌剪枝方法的局部剪枝缺陷,在多个VideoLLMs上实现良好的准确率-效率权衡,在LLaVA-OneVision-7B上剪去75%视觉令牌,保留100.2%原始性能并获1.17倍端到端加速。
中文摘要 AI 辅助
视频大语言模型(VideoLLMs)在视频理解任务中表现出色,但由于长视频中存在大量冗余的时空视觉令牌,其推理成本仍然很高。现有的令牌剪枝方法通过减少冗余令牌来降低成本,但大多数方法依赖于片段级的局部剪枝:将视频划分为独立的片段,在每个片段内独立选择令牌。这种设计可能会不足地保留语义密集的短片段,还会丢弃在局部不显著但从全局来看至关重要的令牌。为解决该问题,本文提出GSTEP(Global Spatio-Temporal Density Pruning,全局时空密度剪枝),这是一种即插即用的剪枝框架,将视频建模为连续的时空信息流。GSTEP通过结合平滑后的中心帧级变化信号得到的连续时间密度,以及帧内空间密度,构建令牌级时空密度,再通过联合平衡信息密度和覆盖范围执行全局令牌采样。在多个VideoLLMs和公共基准上的大量实验表明,GSTEP始终能实现良好的准确率-效率权衡,且在模型架构和评估设置上具有良好的泛化性。在LLaVA-OneVision-7B上,GSTEP剪去了75%的视觉令牌,保留了基准上原始平均性能的100.2%,并实现了1.17倍的端到端加速。
英文摘要
Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number of redundant spatio-temporal visual tokens in long videos. Existing token pruning methods alleviate this cost by reducing redundant tokens, yet most of them rely on segment-level local pruning, where videos are partitioned into isolated segments and tokens are selected independently within each segment. Such designs may under-preserve short but semantically dense segments and discard tokens that appear non-salient locally but remain critical from a global perspective. To address this issue, we propose GSTEP (Global Spatio-Temporal Density Pruning), a plug-and-play pruning framework that models video as a continuous spatio-temporal information flow. GSTEP constructs a token-level spatio-temporal density by combining a continuous temporal density, obtained from a smoothed centered frame-level change signal, with intra-frame spatial density, and then performs global token sampling by jointly balancing information density and coverage. Extensive experiments on multiple VideoLLMs and public benchmarks demonstrate that GSTEP consistently achieves strong accuracy-efficiency trade-offs and generalizes well across model architectures and evaluation settings. On LLaVA-OneVision-7B, GSTEP prunes 75% of visual tokens, preserves up to 100.2% of the original average performance across benchmarks, and achieves a 1.17 end-to-end speedup.