何时检索,何时保持:面向流式视频大语言模型的不确定性感知时间证据分配
When to Retrieve, When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLMs
- IIAU Lab, Dalian University of Technology(大连理工大学IIAU实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对流式视频理解中证据选择平衡时间近因与查询相关性的问题,提出无需训练的WRWS框架,利用轻量级编码器评分与熵代理不确定性,自适应分配语义检索与近因先验,在多个基准上取得竞争力精度并显著降低推理时间。
AI中文摘要:
流式视频理解要求视频大语言模型(Video-LLMs)在因果约束下对连续的视觉流进行推理。随着视觉历史的增长,有限的视觉处理预算要求证据选择在时间近因性与查询相关性之间取得平衡。仅选择最近证据会排除潜在相关的历史证据,而仅进行语义检索在相关性分数模糊时可能会置换有用的近期上下文。我们提出了WRWS(何时检索,何时保持),一个无需训练的框架,用于不确定性自适应的证据分配。一个轻量级的外部视觉-语言编码器对观察到的历史进行查询相关性评分,而一个自适应分配模块使用相似性分布的归一化熵作为检索不确定性的代理。WRWS在相关性线索可靠时倾向于语义检索,在不确定性下增强近因性先验。遵循先检索后编码的流程,WRWS在目标模型视觉编码之前选择证据,使得只有选定的观察结果被昂贵的目标视频大语言模型处理。在四个视频大语言模型家族和多个模型规模上的实验表明,在StreamingBench和OVO-Bench上具有竞争力的准确性。在我们的效率评估中,WRWS将平均视觉到回答时间减少到最先进方法的47.93%。代码将发布。
英文摘要:
Streaming video understanding requires Video Large Language Models (Video-LLMs) to reason over continuous visual streams under causal constraints. As the visual history grows, a bounded visual?processing budget requires evidence selection that balances temporal recency with query relevance. Recent-only selection excludes potentially relevant historical evidence, whereas Semantic-only retrieval can displace useful recent context when relevance scores are ambiguous. We introduce WRWS (When to Retrieve, When to Stay), a training-free framework for uncertainty-adaptive evidence allocation. A lightweight external vision-language encoder scores query relevance across the observed history, while an adaptive allocation module uses the normalized entropy of the similarity distribution as a proxy for retrieval uncertainty. WRWS favors semantic retrieval when relevance cues are reliable and strengthens the recency prior under uncertainty. Following a retrieve-first, encode-later pipeline, WRWS selects evidence before target-model visual encoding, such that only the selected observations are processed by the costly target Video-LLM. Experiments across four Video-LLM families and multiple model scales demonstrate competitive accuracy on StreamingBench and OVO-Bench. In our efficiency evaluation, WRWS reduces average vision-to-answer time to 47.93% of the state-of-the-art method. Code will be released.