arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23265cs.CV

WaveZip:用于视频令牌压缩的小波驱动时空解耦

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

Yuhui Zeng, Wang Chen, Jinfa Huang, Tianyu Xie, Yongdong Luo, Jiayi Ji, Xiawu Zheng, jiebo Luo

首次发表
浏览论文内容

中文总结 AI 辅助

针对长视频理解中视觉令牌计算成本高的问题,提出WaveZip框架,利用离散小波变换分离时空信号,无需特定任务训练,能集成到LVLMs提升推理效率,在长视频理解基准测试中表现优异。

中文摘要 AI 辅助

现有大型视觉语言模型(LVLMs)因视觉令牌的二次计算成本而难以进行长视频理解。近期方法在空间特征域压缩令牌,未分离结构上下文和语义细节。本文提出WaveZip,一个用于高效视频推理的联合信号频域框架。利用离散小波变换(DWT)分离时空信号,通过1D DWT分析查询帧相关性,高频系数由帧间差异门控,共同驱动帧级令牌预算动态分配;2D DWT分解空间特征,高频系数在查询显著区域调制以调节空间重建。WaveZip无需特定任务训练,可无缝集成到现成LVLMs中提高推理效率。在长视频理解基准测试中,WaveZip在10倍极端压缩率下保留99.6%的完整性能,持续优于现有方法。

英文摘要

Existing Large Vision-Language Models (LVLMs) struggle with long-form video understanding due to the quadratic computational cost of visual tokens. While recent efficient methods attempt to compress tokens via hard pruning or uniform merging, they operate strictly in the spatial feature domain, where robust structural context and discriminative semantic details are inherently entangled. In this work, we propose WaveZip, a joint signal-frequency-domain framework for efficient video inference. Driven by the insight that temporal redundancy resides in low-pass approximation scales while spatial saliency strongly correlates with high-frequency components, WaveZip leverages Discrete Wavelet Transforms (DWT) to disentangle these signals. Temporally, it employs 1D DWT to analyze query-frame relevance, and the resulting high-frequency coefficients are further gated by inter-frame differences, with both signals jointly driving the dynamic allocation of a precise frame-level token budget. Spatially, a 2D DWT decomposes features into low-frequency approximations and high-frequency detail components, where the high-frequency coefficients are modulated within query-salient regions to regulate spatial reconstruction. Importantly, WaveZip requires no task-specific training and can be seamlessly integrated into off-the-shelf LVLMs to boost inference efficiency. Extensive experiments on long video understanding benchmarks demonstrate that WaveZip retains 99.6% of the full performance under an extreme 10x compression ratio, consistently outperforming state-of-the-art methods.

发表机构

  • Media Analytics and Computing Lab, Xiamen University(厦门大学媒体分析与计算实验室)
  • Department of Computer Science, University of Rochester(罗切斯特大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑