arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考基于全局码本的视频令牌压缩:一次学习,处处压缩

Rethinking Video Token Compression with a Global Codebook: Learning Once, Compressing Everywhere

Jiayang He, Tianling Xu, Diancheng Kang, Huaide Jiang, Junyan Bai, Shaoming Zheng, Xuan Song

arXiv 2608.01271首次发表:更新:

AI 中文总结

该研究提出即插即用框架ONCE,通过离线学习全局码本实现视频令牌的高效压缩,在保持性能的同时降低推理延迟,提升了视频大语言模型的效率与通用性。

AI 中文摘要

视频大语言模型(Video-LLMs)将视频表示为密集的视觉令牌序列,其长度随输入的时空范围增长。这些令牌常因重复视觉模式产生大量冗余,导致后续语言模型处理出现不必要的计算。现有令牌压缩方法(包括剪枝和合并)在推理时在线执行压缩,对每个输入视频反复产生额外计算,且常依赖特定模型设计,限制了通用性。我们转而通过将高成本压缩过程转为离线模式重新思考该范式,提出了\textbf{ONCE}——一种即插即用的视频令牌压缩框架,引入了离线转在线范式:在视觉特征空间中一次学习频率感知的全局码本,再通过码本查找与聚合用于轻量在线压缩,减少了每个视频的重复计算,且无需特定模型的压缩设计。在多个视频理解基准上针对多种压缩基线开展的大量实验表明,我们的方法实现了出色的准确率-效率权衡,在保持竞争力性能的同时,达到了对比方法中最低的推理延迟。

英文摘要

Video large language models (Video-LLMs) represent videos as dense sequences of visual tokens, whose length grows with the temporal and spatial extent of the input. These tokens often contain substantial redundancy arising from repeated visual patterns, leading to unnecessary computation in the subsequent language-model processing. Existing token compression methods, including pruning and merging, perform compression online during inference, repeatedly incurring additional computation for each input video and often relying on model-specific designs that limit their generality, we instead rethink this paradigm by shifting the costly compression process offline. We propose \textbf{ONCE}, a plug-in video token compression framework that introduces an offline-to-online paradigm: a frequency-aware global codebook is learned once in the visual feature space and reused for lightweight online compression through codebook lookup and aggregation, reducing repeated per-video computation and the need for model-specific compression designs. Extensive experiments across multiple video understanding benchmarks and against diverse compression baselines demonstrate that our approach achieves a strong accuracy-efficiency trade-off, maintaining competitive performance while achieving the lowest inference latency among compared methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑