面向视频多模态大语言模型的视觉令牌编码
Visual Token Coding for Video Multimodal Large Language Models
浏览论文内容
中文总结 AI 辅助
本文提出面向视频多模态大语言模型的视觉令牌编码VTC,新增动态设计得到$VTC_{Dy}$,在Qwen3-VL等模型上实验显示其在低令牌预算下仍保持高性能,且为即插即用设计无需额外调优。
中文摘要 AI 辅助
本文提出了一种面向视频多模态大语言模型(Multimodal Large Language Models, MLLMs)的新型令牌压缩范式,称为视觉令牌编码(Visual Token Coding, VTC)。受HEVC等经典视频编码原理启发,VTC通过预测视频的I帧/P帧并测量其帧级残差来估计令牌冗余度,以此执行结构化压缩。在该基础框架上,我们还为VTC增添了一系列新颖的动态设计,如动态分辨率输入(Dynamic Resolution Input, DyRSO)、动态令牌分配(Dynamic Token Allocation, DyTA)和空间覆盖Top-K(Spatial Coverage Top-K, SC-TopK),并将这种新方法命名为$VTC_{Dy}$。为验证VTC的有效性,我们将其应用于三个MLLMs,并在多个视频理解基准上开展实验。实验结果显示,在令牌预算为50%时,$VTC_{Dy}$在Qwen3-VL上实现了100.1%的平均性能保留率;当令牌预算降至25%时,仍能保留97.8%的平均性能。此外,作为一种即插即用的设计,VTC无需对MLLMs进行额外的令牌编码调优。我们的代码可在此https URL获取。
英文摘要
In this paper, we propose a new token compression paradigm for video Multimodal Large Language Models (MLLMs), termed Visual Token Coding (VTC). Inspired by classical video coding principles, e.g., HEVC, VTC performs structured compression by predicting the I/P frames of a video and measuring their frame-wise residuals to estimate token redundancy. Based on this baseline framework, we also enhance VTC with a set of novel dynamic designs, such as Dynamic Resolution Input (DyRSO), Dynamic Token Allocation (DyTA), and Spatial Coverage Top-K (SC-TopK), and term this new approach $VTC_{Dy}$. To validate VTC, we apply it to three MLLMs and conduct experiments on multiple video understanding benchmarks. The experimental results show that VTC$_{\mathrm{Dy}}$ achieves an average performance retention of 100.1% with a 50% token budget for Qwen3-VL, while still retaining 97.8% of the average performance when the token budget is reduced to 25%. Moreover, as a plug-and-play design, VTC requires no additional tuning of MLLMs for token coding. Our code is available at https://github.com/Msr233/VTC.
发表机构
- Xiamen University(厦门大学)
机构由 AI 辅助整理,请以论文原文为准。