arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MEMO:面向流式视频理解的多层级实体感知记忆

MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding

Yinying Li, Yuqian Fu, Yulin Dai, Jingyu Gong, Tianwen Qian, Xiaoling Wang

arXiv 2609.38900首次发表:更新:

发表机构

East China Normal University; King Abdullah University of Science and Technology(华东师范大学; 阿卜杜拉国王科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MEMO框架,通过多层级实体感知结构化记忆建模流式视频,实现无需训练的即插即用,在多个基准上取得最优性能。

AI 中文摘要

流式视频理解要求模型处理无界视觉流,同时在广阔的时间范围内保留丰富的视觉语义,这对记忆建模提出了根本性挑战。现有方法主要侧重于增加记忆容量,要么将历史信息压缩为固定大小的表示,要么将存储扩展到GPU内存之外。然而,这些方法大多依赖全局或粗粒度的表示,不可避免地丢失细粒度的视觉信息。在这项工作中,我们认为流式视频记忆应显式编码结构化且语义有意义的表示,尤其是在实体层级。为此,我们提出了MEMO,一种通过多层级、实体感知的结构化记忆对流式视频进行建模的新框架。MEMO执行多层级感知,以联合捕捉全局语义、实体动态和空间结构,将流式视频划分为语义连贯的块。每个块被组织成结构化记忆,其中轻量级的全局和实体级表示作为检索索引,而相应的高分辨率视觉内容则单独保留以供按需访问。在推理时,MEMO对结构化记忆执行查询特定的检索,并选择性地召回相关视觉证据以供下游推理使用。值得注意的是,MEMO无需训练,可与现有多模态大语言模型即插即用。在StreamingBench和OVO-Bench上的大量实验表明,MEMO持续改进多个基础模型,并实现了最先进的性能。

英文摘要

Streaming video understanding requires models to process unbounded visual streams while preserving rich visual semantics across vast temporal horizons, posing a fundamental challenge for memory modeling. Existing approaches primarily focus on increasing memory capacity, either by compressing historical information into fixed-size representations or by extending storage beyond GPU memory. However, these methods largely rely on global or coarse-grained representations, inevitably losing fine-grained visual information. In this work, we argue that streaming video memory should explicitly encode structured and semantically meaningful representations, particularly at the entity level. To this end, we propose MEMO, a novel framework that models streaming video through multi-level, entity-aware structured memory. MEMO performs multi-level perception to jointly capture global semantics, entity dynamics, and spatial structures, partitioning streaming video into semantically coherent chunks. Each chunk is organized into a structured memory, where lightweight global and entity-level representations serve as retrieval indices, while the corresponding high-resolution visual content is retained separately for on-demand access. At inference time, MEMO performs query-specific retrieval over the structured memory and selectively recalls relevant visual evidence for downstream reasoning. Notably, MEMO is training-free and plug-and-play with existing multimodal large language models. Extensive experiments on StreamingBench and OVO-Bench demonstrate that MEMO consistently improves multiple base models and achieves state-of-the-art performance.

CommentsAccepted by ACM Multimedia 2026 (ACM MM 2026)

DOI:10.1145/3767308.3835688

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑