集式推理的流式视频令牌压缩
Think in Sets for Streaming Video Token Compression
浏览论文内容
中文总结 AI 辅助
该研究针对流式视频令牌压缩的参考集困境,提出无训练的NovaCov集式压缩器,在保留ReKV 99.6%准确率的同时降低46%的大语言模型预填充延迟,性能优于现有无训练方法。
中文摘要 AI 辅助
流式视频大语言模型(Streaming VideoLLMs)按因果关系处理帧,视觉令牌持续增长,因此压缩对控制预填充延迟和内存至关重要。现有无训练方法独立对令牌排序,忽略保留令牌间的边际增益交互。我们认为,流式视频令牌压缩应被表述为集选择,每个候选令牌的价值由其在已保留令牌之外新增的内容决定。与现有为离线任务设计的集式方法不同,流式处理要求逐帧做出因果剪枝决策,因此建模跨帧交互需要显式的历史参考,这就产生了参考集困境:参考集必须充分代表先前已传递的内容,同时又要受限于实时推理的约束。我们提出NovaCov,据我们所知,这是首个专为流式视频设计的无训练、即插即用的集式令牌压缩器。NovaCov维护一个容量受限、按近期性加权的历史参考库,并优化一个双分支次模覆盖目标,该目标在保留当前帧代表性内容的同时,优先处理历史中未充分覆盖的信息。两个分支均为设施定位函数,因此贪心选择保留了经典的(1-1/e)近似保证。在流式和离线基准测试中,NovaCov的性能优于现有无训练压缩方法,在保留ReKV 99.6%准确率的同时,将大语言模型预填充延迟降低了46%。
英文摘要
Streaming VideoLLMs process frames causally while visual tokens grow continuously, making compression essential for controlling prefilling latency and memory. Existing training-free methods independently rank tokens, ignoring marginal-gain interactions among retained tokens. We argue that streaming video token compression should instead be formulated as set selection, where each candidate is valued by what it adds beyond the tokens already retained. Unlike existing set-wise methods designed for offline tasks, streaming makes causal, frame-by-frame pruning decisions, so modeling cross-frame interactions requires an explicit historical reference. This creates a reference-set dilemma: the reference must adequately represent previously conveyed content while remaining bounded for real-time inference. We introduce NovaCov, to our knowledge the first training-free, plug-and-play set-wise token compressor designed for streaming video. NovaCov maintains a capacity-bounded, recency-weighted Historical Reference Bank and optimizes a dual-branch submodular coverage objective that preserves representative current-frame content while prioritizing information insufficiently covered by history. Both branches are facility-location functions, so greedy selection retains the classical (1-1/e) approximation guarantee. Across streaming and offline benchmarks, NovaCov outperforms existing training-free compression methods, retaining 99.6% of ReKV accuracy while reducing LLM prefilling latency by 46%.