发表机构
NJU; PJLAB; SJTU; USTC; CAS; CUHK; PKU; THU; FDU; ZJU(南京大学; 上海人工智能实验室; 上海交通大学; 中国科学技术大学; 中国科学院; 香港中文大学; 北京大学; 清华大学; 复旦大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出OneStreamer,通过主动生成统一流式视频交互中的感知、记忆与响应,其PHCM和PSTL机制在8个基准上取得最优性能,并构建了百万级数据集OneStreamer-1M。
AI 中文摘要
流式视频大语言模型(LLM)必须在证据与未来任务的相关性尚不明确之前保留证据,并在获得充分证据时做出响应。挑战在于形成可复用的事实记忆,同时不损害实时感知能力。我们提出OneStreamer,通过共享的主动生成过程,联合学习与查询无关的证据记录和任务响应。其主动分层字幕记忆(PHCM)生成带有时间锚定的局部细节字幕以及已完成事件的摘要。在训练期间,流式字幕目标监督对观测到的视频前缀的解释。在推理时,模型生成的记录补充近期视觉窗口,提供可复用的事实上下文,而无需重新访问历史视觉特征。主动状态转换学习(PSTL)通过在所有输出锚点保留监督并选择具有代表性的状态变化和状态持续令牌,减少了重复等待状态的主导地位。我们进一步开发了一个流式数据合成流水线,使输出内容和时间与可用证据对齐。将生成的流式字幕和问答与清理后的开源数据相结合,得到OneStreamer-1M,这是一个覆盖广泛任务的流式视频交互数据集,包含超过一百万条记录。我们的4B模型在全部八个评估的流式视频理解基准中,与对比方法相比取得了最佳结果。消融实验表明,保留生成的字幕可改善历史问答,且不降低实时感知能力。PSTL在仅监督27.5%的标注状态令牌的情况下,也优于密集状态监督。综合来看,这些结果支持主动生成作为连接流式视频交互中感知、记忆形成和及时响应的共享学习接口。
英文摘要
Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.
Comments29 pages, 12 figures, 20 tables. Project page: https://mcg-nju.github.io/OneStreamer