CoFiE:用于高效流视频理解的由粗到细证据选择框架
CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding
浏览论文内容
中文总结 AI 辅助
CoFiE 是一种由粗到细证据选择框架,通过解耦证据选择阶段减少流视频理解的延迟,在多个基准上实现最优准确率-效率权衡,性能优于现有模型
中文摘要 AI 辅助
流视频理解要求视觉语言模型(VLLMs)在严格的延迟约束下处理不断增长的视频流并回答用户问题。现有方法通过 token 剪枝和记忆库方案提升效率,但主要在视觉编码后减少视觉 token,仅下游 token 剪枝无法大幅降低端到端延迟,因为昂贵的帧编码成本已产生。本文提出 CoFiE,一种由粗到细证据选择框架,将证据选择解耦为视觉编码器前的与查询无关的粗过滤阶段,以及 LLM 预填充期间的与查询相关的细优化阶段。CoFiE 引入新颖性引导帧过滤以保留视觉上有辨识度的候选帧,以及查询特定证据优化以选择与用户查询最相关的帧。该设计在帧编码前去除大量冗余,同时在语义信息可用后保留查询特定优化。实验表明,CoFiE 在多个视频理解基准上实现了新的最优准确率-效率权衡,在 StreamingBench 上达到 78.86% 的准确率,在 OvO-Bench 上达到 68.72%,比现有方法提升最高达 3.15%。即使在证据帧过滤比例高达 80% 的情况下,CoFiE 仍优于强大的开源多模态模型,同时将端到端推理延迟降低最高达 2.54 倍。
英文摘要
Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and answer user questions under tight latency constraints. Existing methods improve efficiency through token pruning and memory-bank schemes, but mainly reduce visual tokens after visual encoding. Consequently, downstream token pruning alone cannot substantially reduce end-to-end latency because the expensive frame encoding cost has already been incurred. We propose CoFiE, a Coarse-to-Fine Evidence Selection framework that decouples evidence selection into a coarse, query-agnostic filtering stage before the vision encoder and a fine, query-specific refinement stage during LLM prefill. CoFiE introduces Novelty-Guided Frame Filtering to retain visually distinctive candidate frames and Query-Specific Evidence Refinement to select the frames most relevant to the user query. This design removes substantial redundancy before frame encoding while preserving query-specific refinement once semantic information becomes available. Experiments show that CoFiE establishes a new state-of-the-art accuracy-efficiency trade-off across multiple video understanding benchmarks, reaching 78.86% accuracy on StreamingBench and 68.72% on OvO-Bench, with improvements of up to 3.15% over prior methods. Even with up to 80% evidence-frame filtering, CoFiE outperforms strong open-source multimodal models while improving end-to-end inference latency by up to 2.54 times.
发表机构
- Harbin Institute of Technology(哈尔滨工业大学)
- State Key Laboratory of Smart Farm Technologies and Systems(智慧农场技术与系统国家重点实验室)
- Pengcheng Laboratory(鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。