基于最优传输的视觉信息聚合用于视频语言模型的令牌压缩
Aggregating Visual Information with Optimal Transport for VideoLM Token Compression
中文总结 AI 辅助
该研究针对视频语言模型的视觉令牌冗余问题,提出AVIOT方法,将视频令牌压缩转化为最优传输问题,沿任务和空间轴调整表示构造,在多基准上实现了与未压缩基线相当或更优的性能,且高压缩率下表现稳定。
中文摘要 AI 辅助
视频语言模型将视频处理为具有大量表示冗余的密集视觉令牌序列,因此压缩这些序列对于减轻语言模型解码的视觉令牌负担至关重要。核心挑战在于在压缩过程中保留跨帧分散的视觉信息。为此,我们提出了基于最优传输的视觉信息聚合(AVIOT),其将视频令牌压缩建模为将帧观测的密集经验测度传输到紧凑目标测度上。所得的源到目标耦合为每个目标支撑集诱导出源观测的分布,直接指定压缩视频表示的构造方式。我们进一步沿任务和空间轴调整该构造:问题条件会调制源帧与目标支撑集之间的传输代价,同时影响分配给每个时间片段的支撑集数量,从而将表示容量导向与问题相关的内容。在多个空间粒度上,AVIOT计算特定区域的时间传输计划并自适应融合其产生的表示,使得同一紧凑表示内的不同区域可从不同时刻获取信息。在不同压缩率下的评估表明,AVIOT在多个视频理解基准上与未压缩基线相当或优于基线,且在更高压缩率下仍保持较强性能。
英文摘要
Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing the visual-token burden on language-model decoding. The central challenge is to preserve visual information dispersed across frames under such compression. To this end, we introduce Aggregating Visual Information with Optimal Transport (AVIOT), which casts video token compression as transporting a dense empirical measure of frame observations onto a compact target measure. The resulting source-to-target coupling induces a distribution over source observations for each target support, directly specifying how the compressed video representation is constructed. We further adapt this construction along task and spatial axes. Question conditioning modulates the transport cost between source frames and target supports, while influencing how many supports are allocated to each temporal segment, thereby directing representation capacity toward question-relevant content. At multiple spatial granularities, AVIOT computes region-specific temporal transport plans and adaptively fuses the representations they yield, allowing different regions within the same compact representation to draw from different moments. Evaluations across varying compression ratios show that AVIOT matches or outperforms the uncompressed baseline on multiple video-understanding benchmarks while retaining strong performance at higher compression ratios.