arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CRAFT:面向视觉语言模型的基于视频令牌递归自适应融合的压缩方法

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

Yu Chen, Xiaohong Li, Xiaole Wang, Jianjin Zhang, Jun Sun, Yafeng Deng

arXiv 2608.01644首次发表:更新:

AI 中文总结

针对视觉语言模型视频令牌压缩效率与适应性的权衡问题,提出CRAFT方法,通过解耦令牌选择与融合实现高效压缩,在8倍压缩时保留97%主干准确率且性能优于现有方法。

AI 中文摘要

在视频理解领域,视觉语言模型(VLMs)需处理海量视觉令牌,这使得预填充阶段的计算与内存成本急剧上升。这类视觉序列在时空维度存在高度冗余,但高压缩率往往伴随关键细节的丢失。现有的令牌压缩方法要么采用无训练的启发式压缩,内容适应性有限;要么引入额外模块,需昂贵的对齐训练,导致效率与适应性之间的权衡问题未得到解决。为缓解这一局限,我们提出CRAFT:基于视频令牌递归自适应融合的压缩方法。CRAFT通过将无参数的令牌选择与可学习的令牌融合解耦,递归合并令牌:全局相似度决定需合并的令牌,而位置感知加权模块与内容自适应通道门控则学习如何融合这些令牌。整个压缩流程与查询无关。由于每个保留的令牌都是原始令牌的线性组合,CRAFT保留了其真实时空坐标,并与预训练语言模型的输入分布保持对齐。在多个代表性视频基准上的实验表明,CRAFT始终优于现有的最先进令牌压缩方法。在约8倍压缩率下,它保留了主干网络约97%的平均准确率,并展现出显著的效率提升。

英文摘要

In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply. Such visual sequences are highly redundant along the spatio-temporal dimension, yet a high compression ratio is often accompanied by the loss of critical details. Existing token-compression methods either employ heuristic, training-free compression with limited content adaptivity or introduce additional modules that require expensive alignment training, leaving the trade-off between efficiency and adaptivity unresolved. To alleviate this limitation, we propose CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens. CRAFT recursively merges tokens by decoupling parameter-free token selection from learnable token fusion: global similarity determines which tokens to merge, while a position-aware weighting module and a content-adaptive channel-wise gate learn how to fuse them. The whole compression pipeline is query-agnostic. Because every retained token is a linear combination of the original tokens, CRAFT preserves their true spatio-temporal coordinates and stays aligned with the pre-trained language model's input distribution. Experiments on multiple representative video benchmarks show that CRAFT consistently outperforms prior state-of-the-art token-compression methods. At about $8\times$ compression, it retains roughly $97\%$ of the backbone's average accuracy and shows significant efficiency improvement.

Comments11 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑