用于流视频理解的动态轮辐式记忆
Dynamic Hub-and-Spoke Memory for Streaming Video Understanding
浏览论文内容
中文总结 AI 辅助
针对流视频理解需同时紧凑记忆长程历史与高效检索相关证据的挑战,提出动态轮辐式记忆(D-HSM)框架,结合结构化文本记忆与近期视觉标记,在基准上提升VLM性能并优于现有基线。
中文摘要 AI 辅助
流视频理解需要针对不断增长的连续视觉流在任意时刻回答问题,核心挑战是在紧凑记忆长范围历史的同时有效检索与问题相关的证据。我们提出动态轮辐式记忆(Dynamic Hub-and-Spoke Memory, D-HSM),这是一种无需训练的框架,它将遥远历史表示为结构化文本记忆,同时保留近期帧作为视觉标记以实现细粒度感知。具体而言,D-HSM将选定的历史视频片段转换为带类型的文本观测值,并存储在以实体为中心的轮辐式记忆中,其中实体作为中心,相关证据作为轮辐。在回答问题时,D-HSM动态检索一个紧凑的问题感知记忆子集,通过轮辐式链接扩展该子集,并将其与近期视觉窗口结合,用于冻结视觉语言模型(frozen-VLM)的答案预测。在流视频和长视频基准上的大量实验表明,D-HSM能持续且显著地提升视觉语言模型(VLM)骨干的性能,并且优于其他最先进的在线和离线视频理解基线。
英文摘要
Streaming video understanding requires answering questions at arbitrary times over a continuously growing visual stream. The central challenge is to compactly remember long-range history while effectively retrieving question-relevant evidence. We propose Dynamic Hub-and-Spoke Memory (D-HSM), a training-free framework that represents distant history as structured textual memory while preserving the recent frames as visual tokens for fine-grained perception. Specifically, D-HSM turns selected historical video chunks into typed textual observations and stores them in an entity-centered hub-and-spoke memory, with entities as hubs and related evidence as spokes. When answering a question, D-HSM dynamically retrieves a compact question-aware memory subset, expands it through hub-and-spoke links, and combines it with the recent visual window for frozen-VLM answer prediction. Extensive experiments on both streaming and long video benchmarks show that D-HSM consistently and substantially improves VLM backbones and outperforms other state-of-the-art online and offline video understanding baselines.
发表机构
- Northeastern University(东北大学)
- University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)
- Tulane University(杜兰大学)
- University of Virginia(弗吉尼亚大学)
- Adobe Research(奥多比研究院)
机构由 AI 辅助整理,请以论文原文为准。