用于视频-语言模型高效推理的自适应两阶段视觉令牌剪枝
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models
浏览论文内容
中文总结 AI 辅助
针对视频-语言模型推理延迟高的问题,提出两阶段自适应令牌剪枝策略,先剪去视频冗余帧,再根据视频内容自适应剪去保留帧中的冗余令牌,无需额外训练微调,在10%令牌保留率下提升准确率7%并降低95%计算量。
中文摘要 AI 辅助
视觉-语言模型在图像和视频理解方面表现出色,但由于每张图像需处理数千个令牌,导致推理延迟较高,限制了其在资源受限的边缘设备和实时监控应用中的部署。在视频处理中,需同时分析多帧,这一挑战进一步加剧。现有的令牌减少技术主要针对单图像输入开发,因此未考虑视频序列中存在的时间和帧间冗余。此外,这些方法通常对所有输入应用固定的统一剪枝比例,这并非最优,因为不同视频间的冗余程度差异显著,需要依赖内容的剪枝级别来保留关键信息。为解决这些局限,我们提出一种专门针对视频处理的两阶段自适应令牌剪枝策略。第一阶段,剪去冗余帧;第二阶段,在保留的帧内应用令牌级剪枝。关键在于,第二阶段的剪枝比例根据每个视频的内容自适应确定,具体通过分析令牌嵌入的相关结构来量化冗余,以此确定比例。重要的是,我们的方法完全是事后的,无需额外训练或微调,同时取得了显著的经验增益;值得注意的是,在令牌保留率为10%时,它在视频字幕基准测试上将准确率提高了7%,同时将计算量(TFLOPs)降低了95%。
英文摘要
Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands of tokens per image, limiting their deployment on resource-constrained edge devices and in real-time surveillance applications. This challenge is further amplified in video processing, where multiple frames must be analyzed simultaneously. Existing token reduction techniques are largely developed for single-image inputs and therefore fail to account for the temporal and inter-frame redundancies present in video sequences. In addition, these methods generally rely on a fixed, uniform pruning ratio applied across all inputs, which is suboptimal because the degree of redundancy can vary significantly between different videos, necessitating content-dependent pruning levels to preserve critical information. To address these limitations, we propose a two-stage adaptive token pruning strategy specifically designed for video processing. In the first stage, we prune out the redundant frames, and in the second stage, token-level pruning is applied within the retained frames. Crucially, the pruning ratio in the second stage is determined adaptively based on the content of each video. This is achieved by analyzing the correlation structure of token embeddings to quantify redundancy, which is used to determine the ratio. Importantly, our method is entirely post-hoc and requires no additional training or fine-tuning, while achieving strong empirical gains; notably, it improves accuracy by +7\% on a video captioning benchmark at 10\% token retention, while reducing computation TFLOPs by 95\%.