arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Watch-Think-Interact: 通过强化学习引导长时程多轮流式视频推理

Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning

Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao, Bangwei Liu, Kunquan Li, Yuxiang Shen, Bangyan Li, Xuejiao Wang, Changbo Wang, Gaoqi He

arXiv 2609.37035首次发表:更新:

发表机构

East China Normal University; University of Science and Technology of China; Xiamen University(华东师范大学; 中国科学技术大学; 厦门大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对流式视频多轮推理中在线状态遗漏视觉细节的问题,提出WTI闭环框架,结合自然语言记忆与选择性回忆,并构建WTI-82K数据集及Stream-GDPO强化学习算法,在StreamingBench和OVO-Bench上取得最优性能。

AI 中文摘要

流式视频辅助要求模型在固定上下文预算下,根据观察到的前缀回答异步问题。现有方法对响应时序进行建模或压缩历史,但在未来问题已知之前形成的在线状态可能会遗漏视觉细节,而这些细节在后续问题揭示其相关性之前可能未被注意到;仅保留的状态无法恢复这些细节。我们提出了Watch-Think-Interact (WTI),一个用于多问题流式视频推理的闭环框架。WTI维护带有源视频时间范围标签的紧凑自然语言记忆条目;这些条目在足够时支持直接推理,否则锚定对更精细视觉证据的选择性回忆。对于每个问题,WTI在当前上下文和记忆足够时回答,当所需证据尚未出现时继续观看,或回忆相关的过去区间并在整合返回的块后再次决定,而不重放完整的观察历史。为了训练这种行为,我们构建了WTI-82K,包含4,812条因果对齐轨迹中的82,335个定时问题,并开发了Stream-GDPO,利用轨迹级反馈优化完整的多问题流式滚动,用于响应时序、源视频回忆和记忆更新。WTI在比较的开源流式基线中实现了最先进的聚合性能,在StreamingBench上达到83.3%,在OVO-Bench上达到73.6%的加权总体准确率。

英文摘要

Streaming video assistance requires models to answer asynchronous questions from an observed prefix under a fixed context budget. Existing approaches model response timing or compress history, but an online state formed before future questions are known can omit visual details before later questions reveal their relevance; the retained state alone cannot recover them. We introduce Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning. WTI maintains compact natural-language memory entries tagged with source-video time ranges; these entries support direct reasoning when sufficient and otherwise anchor selective recall of finer visual evidence. For each question, WTI answers when current context and memory suffice, continues watching when required evidence has not appeared, or recalls a relevant past interval and decides again after incorporating the returned chunks, without replaying the full observed history. To train this behavior, we construct WTI-82K, comprising 82,335 timed questions across 4,812 causally aligned trajectories, and develop Stream-GDPO to optimize complete multi-question streaming rollouts using trajectory-level feedback for response timing, source-video recall, and memory updates. WTI achieves state-of-the-art aggregate performance among the compared open-source streaming baselines, reaching 83.3% on StreamingBench and 73.6% weighted overall accuracy on OVO-Bench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑