发表机构
National Taiwan University; NVIDIA(国立台湾大学; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VETO通过双轴压缩(帧内语义合并与帧间冗余合并)优化视觉令牌,突破长视频VLM推理效率瓶颈,实现高达45%加速并保持或提升精度,在极端令牌预算下优于现有方法。
AI 中文摘要
使用视觉语言模型(VLM)处理长视频时,视觉令牌的二次方成本成为瓶颈,使得长序列推理代价高昂。虽然单轴压缩方法能缓解这一问题,但由于它们独立处理空间和时间冗余,因此遇到了硬性的效率下限。我们提出VETO(视频高效令牌优化,Video Efficient Token Optimization for Vision-Language Models),一种无需训练的即插即用模块,通过双轴压缩消除这一瓶颈:(i)帧内压缩器,利用基于最优传输启发的匹配,合并每帧内语义相似的令牌;(ii)帧间压缩器,识别并合并时间上冗余的帧。关键设计思路在于层次化排序:通过先压缩空间维度,VETO大幅降低了后续全局时间匹配的成本,绕过了单轴方法的效率墙,且其优势随着现代全融合注意力基础设施的发展而增长。实验上,VETO在保持或提升精度的同时,实现了高达45%的推理加速(例如在LLaVA-OneVision-7B上)。在极端令牌匮乏(10%预算)下,VETO以55.7%的准确率优于VFlowOpt(54.9%)、VisionZip(52.6%)和FastV(47.9%)。我们展示了其在LLaVA-OneVision、InternVL-2.5和LongVA上的普适性,所有情况下零样本准确率均保持或提升。
英文摘要
Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We present VETO (Video Efficient Token Optimization for Vision-Language Models), a training-optional plug-in that eliminates this bottleneck through dual-axis compression: (i) an intra-frame compressor that merges semantically similar tokens within each frame via optimal-transport inspired matching, and (ii) an inter-frame compressor that identifies and merges temporally redundant frames. The key design insight is hierarchical ordering: by first compressing spatial dimensions, VETO drastically reduces the cost of subsequent global temporal matching, bypassing the efficiency wall of single-axis approaches, with an advantage that grows with modern fully-fused attention infrastructure. Empirically, VETO achieves up to 45% faster inference (e.g., on LLaVA-OneVision-7B) while preserving or improving accuracy. Under extreme token starvation (10% budget), VETO outperforms VFlowOpt (54.9%), VisionZip (52.6%), and FastV (47.9%) with 55.7% accuracy. We demonstrate universal applicability across LLaVA-OneVision, InternVL-2.5, and LongVA, with zero-shot accuracy preserved or improved in all cases.