发表机构
Central South University; SenseTime Research and Tetras.AI(中南大学; 商汤科技研究与Tetras.AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SimpleCluster,一种无需训练的简单基线,通过位置感知跨帧聚类压缩视频视觉令牌,在多个基准上达到与复杂方法相当或更优的性能,并揭示特征空间保留是关键。
AI 中文摘要
视频大语言模型(Video LLMs)在视频理解方面取得了显著进展,但其推理效率受到长视频产生的大量视觉令牌的制约。近年来,视频令牌压缩方法越来越多地引入复杂的令牌选择、剪枝和合并策略。这引发了一个基本问题:仅通过保留视觉表示中编码的结构,可以获得多少压缩性能?我们通过SimpleCluster来研究这个问题,这是一个简单且无需训练的基线,它在视觉特征空间中执行位置感知的跨帧聚类,并使用每个聚类的原始视觉特征的均值来表示该聚类。在四个视频理解基准和三个代表性视频大语言模型上的大量实验表明,SimpleCluster在各种令牌保留比率下均能达到与近期压缩方法相当或更优的性能,在极低保留率(例如1%)下尤其表现出强大的鲁棒性。为了理解这一行为,我们从局部逼近保真度和全局覆盖度两个方面分析了不同压缩方法保留的特征空间。结果表明,更强的下游性能始终与对原始视觉特征分布的更好保留相关,尤其是其全局覆盖度。这些发现凸显了在高度受限的令牌预算下,特征空间保留是视频令牌压缩的一个重要考虑因素。我们的代码可在以下网址获取:此 https URL。
英文摘要
Video Large Language Models (Video LLMs) have achieved remarkable progress in video understanding, but their inference efficiency is constrained by the large number of visual tokens produced by long videos. Recent video token compression methods increasingly introduce sophisticated strategies for token selection, pruning, and merging. This raises a fundamental question: how much of compression performance can be obtained by simply preserving the structure encoded in the visual representations? We investigate this question with SimpleCluster, a simple and training-free baseline that performs position-aware cross-frame clustering in the visual feature space and represents each cluster using the mean of its original visual features. Extensive experiments across four video understanding benchmarks and three representative Video LLMs show that SimpleCluster achieves competitive or superior performance over recent compression methods across a wide range of token retention ratios, with particularly strong robustness under extremely low retention rates (e.g., 1%). To understand this behavior, we analyze the feature space preserved by different compression methods in terms of local approximation fidelity and global coverage. The results show that stronger downstream performance is consistently associated with better preservation of the original visual feature distribution, especially its global coverage. These findings highlight feature-space preservation as an important consideration for video token compression under highly constrained token budgets. Our code is available at https://github.com/xiaozhang79/SimpleCluster.