arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28312cs.CVcs.AI

ObjectStream:将潜在对象作为记忆锚点用于流视频理解

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

Mingkang Dong, Muxin Pu, Jie Li, Bohan Guo, Songruo Chen, Bin Ren, Xu Zheng, Chen Zhao, Tianwen Qian, Mohamed Elhoseiny, Yuqian Fu

首次发表
浏览论文内容

中文总结 AI 辅助

提出无需训练的 ObjectStream 框架,以潜在对象为记忆锚点组织流视频证据,使 Video-LLMs 无需改动即可推理对象相关内容,在多基准上实现性能提升且内存与推理效率显著优化。

中文摘要 AI 辅助

流视频理解要求模型在未来问题已知前持续保留有用的视觉证据。现有方法主要依据 token 重要性、时间冗余或片段级相关性管理不断增长的视觉上下文,但很少围绕随时间持续存在和演化的对象组织证据。因此,本文提出 ObjectStream,这是一个无需训练的框架,将潜在对象作为流视频理解的记忆锚点。ObjectStream 直接从冻结的 Video-LLM 表示中诱导出空间一致的潜在对象,将它们跨帧链接为持久锚点,并在有限内存预算下维护其历史,无需外部对象检测器或分割模型。基于这些锚点,ObjectStream 保留三种互补形式的证据:持久对象历史、瞬态对象变化和近期视觉上下文。该设计使现有视频大语言模型(Video-LLMs)能够对对象身份、交互和状态变化进行推理,同时保持底层模型不变。在在线流和离线长视频基准上的大量实验证明了其有效性和效率:在线流评估中,ObjectStream 使 Qwen2.5-VL-7B 在 OVO-Bench 实时视觉感知任务上提升 10.0 个百分点,同时将峰值 GPU 内存和首帧推理时间(TTFT)降低约 50%;在离线长视频基准上,它在丢弃 82.5% 的视觉 token 的情况下超过全 token 基线。这些结果表明,潜在对象是紧凑流视频记忆的实用且有效的组织原则。

英文摘要

Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches primarily manage the growing visual context according to token importance, temporal redundancy, or segment-level relevance, but rarely organize evidence around objects that persist and evolve over time. Thus, in this paper, we introduce ObjectStream, a training-free framework that treats latent objects as memory anchors for streaming video understanding. ObjectStream induces spatially coherent latent objects directly from frozen Video-LLM representations, links them across frames into persistent anchors, and maintains their histories under a bounded memory budget, without requiring external object detectors or segmentation models. Built on these anchors, ObjectStream preserves three complementary forms of evidence: persistent object histories, transient object changes, and recent visual context. This design enables existing Video Large Language Models (Video-LLMs) to reason over object identities, interactions, and state changes while leaving the underlying model unchanged. Extensive experiments on online streaming and offline long-video benchmarks demonstrate both effectiveness and efficiency. In online streaming evaluation, ObjectStream improves Qwen2.5-VL-7B by 10.0 points on OVO-Bench Real-Time Visual Perception, while reducing peak GPU mem-ory and TTFT by approximately 50%. On offline long-video benchmarks, it surpasses the full-token baseline while discarding 82.5% of visual tokens. These results highlight latent objects as a practical and effective organizing principle for compact streaming video memory.

补充信息

↑