arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04225cs.CV

FlashGaze:用于高效视频理解的无训练多尺度补丁剪枝

FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding

Ziye Zhu, Yanghao Zhou, Lixing Tan, Jialiang Kang, Shuxuan Li, Xiao Yang

首次发表
浏览论文内容

中文总结 AI 辅助

FlashGaze提出无训练的多尺度补丁剪枝方法,在ViT编码前利用像素差异和四叉树动态规划减少时空冗余,在保持98%精度的同时实现最高5.4倍编码加速和17倍预填充加速,并降低1.8倍内存使用。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)在视频理解方面展现出强大性能,但高效处理长时、高分辨率视频仍具挑战性。此类视频通常包含大量时空冗余,处理冗余视觉标记会带来可避免的计算开销。许多现有方法在视觉变换器(ViT)编码期间或之后剪枝视觉标记,导致大量编码成本未被解决。一些方法在编码前剪枝补丁,但依赖学习到的辅助网络进行补丁选择,增加了额外的训练和推理开销。为解决这些局限,我们提出FlashGaze,一种无训练方法,在ViT编码前减少时空冗余,且不引入辅助网络。FlashGaze利用像素空间差异作为信息损失的代理,并采用四叉树动态规划在固定预算下联合优化补丁丢弃、合并和保留。在两个MLLM骨干网络上的多个基准实验表明,该方法在基本保持精度的同时实现了显著的效率提升。在Qwen3-VL-8B上,FlashGaze在LongVideoBench上保留了全输入基线精度的98%,同时ViT编码和MLLM预填充分别实现了高达5.4倍和17倍的加速,并将峰值GPU内存使用量降低了1.8倍。这些效率提升使模型能够在相同GPU硬件上处理更多帧和更高分辨率的视频,解锁了以往无法企及规模的视频理解。

英文摘要

Multimodal Large Language Models (MLLMs) have demonstrated strong performance in video understanding, yet efficiently processing long, high-resolution videos remains challenging. Such videos often contain substantial spatiotemporal redundancy, and processing redundant visual tokens can incur avoidable computational overhead. Many existing methods prune visual tokens during or after vision transformer (ViT) encoding, leaving much of the encoding cost unaddressed. Some approaches prune patches before encoding but rely on learned auxiliary networks for patch selection, incurring additional training and inference overhead. To address these limitations, we propose FlashGaze, a training-free method that reduces spatiotemporal redundancy before ViT encoding without introducing auxiliary networks. FlashGaze uses pixel-space differences as a proxy for information loss and employs Quadtree Dynamic Programming to jointly optimize patch dropping, merging, and keeping under a fixed budget. Experiments on two MLLM backbones across multiple benchmarks demonstrate substantial efficiency gains while largely preserving accuracy. On Qwen3-VL-8B, FlashGaze retains 98% of the full-input baseline accuracy on LongVideoBench while achieving up to 5.4x and 17x speedups in ViT encoding and MLLM prefill, respectively, and reducing peak GPU memory usage by a factor of 1.8. These efficiency gains enable the model to process videos with more frames and higher resolutions on the same GPU hardware, unlocking video understanding at scales previously out of reach.

发表机构

  • Peking University(北京大学)
  • Canva Research(Canva 研究院)
  • Sun Yat-sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑