arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00291cs.CV

StreamScout:学习何时深入探究以实现流视频理解

StreamScout: Learning When to Look Deeper for Streaming Video Understanding

Ce Zhang, Jing Bi, Jinxi He, Jianshu Zhang, Jingyang Lin, Yunzhong Xiao, Minghao Fu, Yaqi Xie, Zhentao Xie, Weicong Chen, Katia Sycara, Ming Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

StreamScout是自适应流视频理解框架,通过分级视觉视图与优化的停止策略提升性能,在多基准测试中优于现有方法,降低推理成本与令牌消耗。

中文摘要 AI 辅助

流视频理解需要回答在无界视频流的任意时刻到达的问题。现有系统主要关注在有限内存中保留什么内容,然而针对每个查询都采用相同的固定成本程序访问该内存,尽管所需证据存在巨大差异。我们认为,针对每个查询决定访问内存的深度与决定内存应存储什么内容同等重要。为此,我们引入StreamScout,这是一种自适应推理框架,随着流的展开仅在上下文中维护轻量级文本时间线。在查询时,StreamScout逐步用多达三种信息性递增的视觉视图增强时间线:近期帧的扫视、过去流的均匀回溯以及查询显著检索。在每个阶段,如果可用证据足够,模型会立即回答;否则,它会升级到下一个视图。为改进这种停止或升级策略,我们在辅助集上探测级联,并将模型的经验能力边界提炼为轻量级LoRA适配的监督,得到StreamScout-S。我们进一步通过强化学习优化策略,允许模型探索超越提炼决策模仿的停止行为,得到StreamScout-R。在三个主干网络和三个流基准测试中,StreamScout及其变体始终优于现有流方法,同时大幅降低推理成本和令牌消耗;例如在OVO-Bench上,StreamScout-S将Qwen3-VL-8B的性能提升14.65个百分点,且令牌使用量比均匀采样少59%,平均回答时间为1.04秒。

英文摘要

Streaming video understanding requires answering questions that arrive at arbitrary moments over an unbounded video stream. Existing systems primarily focus on what to retain in a bounded memory, yet access that memory using the same fixed-cost procedure for every query, despite substantial variation in the evidence required. We argue that deciding how deeply to access memory for each query is as important as deciding what the memory should store. To this end, we introduce StreamScout, an adaptive inference framework that maintains only a lightweight textual timeline in context as the stream unfolds. At query time, StreamScout progressively augments the timeline with up to three increasingly informative visual views: a glance at recent frames, a uniform look-back over the past stream, and query-salient retrieval. At each stage, the model answers immediately if the available evidence is sufficient; otherwise, it escalates to the next view. To improve this stop-or-escalate policy, we probe the cascade on an auxiliary set and distill the model's empirical competence boundary into supervision for a lightweight LoRA adaptation, yielding StreamScout-S. We further refine the policy through reinforcement learning, allowing the model to explore stopping behaviors beyond imitation of the distilled decisions, yielding StreamScout-R. Across three backbones and three streaming benchmarks, StreamScout and its variants consistently outperform prior streaming methods while substantially reducing inference cost and token consumption; on OVO-Bench, for instance, StreamScout-S improves Qwen3-VL-8B by 14.65 points while using 59% fewer tokens than uniform sampling and answering in 1.04 s on average.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • University of Rochester(罗切斯特大学)
  • Northwestern University(西北大学)
  • University of California, San Diego(加利福尼亚大学圣迭戈分校)
  • TikTok(字节跳动旗下TikTok)

机构由 AI 辅助整理,请以论文原文为准。

↑