arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

保持简洁:用于超长视频理解的多键情景记忆检索

Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding

Yeeun Choi, Youngbeom Yoo, Joon-Young Lee, Hyolim Kang, Seon Joo Kim

arXiv 2608.07663首次发表:更新:

发表机构

Yonsei University; Adobe Research(延世大学; 奥多比研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对超长视频理解的两阶段范式需求,提出MERIT框架,通过多键表示与按需时间扩展实现精准检索,在三个长视频基准数据集上取得最优性能。

AI 中文摘要

当视频时长从数小时扩展至数天时,当前多模态大语言模型(MLLM)直接对其进行端到端处理变得不切实际。这种超长场景需要两阶段范式:与查询无关的记忆构建,以及基于检索的推理。现有研究致力于在记忆构建阶段预建模视频的高层关系,却在构建时未知下游查询。我们转而在记忆构建阶段优先保证高召回率的可检索性,将查询特定的高层关系组合推迟至推理阶段完成。为此,我们提出MERIT(Multi-key Episodic Retrieval with Inference-time Temporal expansion,即带推理时时间扩展的多键情景检索),这是一个简单却有效的超长视频理解智能体框架。首先,我们构建情景多键表示,通过简单的键匹配机制实现对细粒度记忆的精准检索;其次,我们引入邻居过滤机制,仅在推理阶段围绕检索到的片段扩展时间范围,以此捕获更广泛的语义上下文,同时避免全局记忆构建带来的巨大计算开销。通过结合简单的键匹配与这种按需时间扩展,MERIT在三个长视频基准数据集EgoLifeQA、LVBench和Video-MME(Long)上取得了最优性能。

英文摘要

When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Models (MLLMs). This ultra-long setting necessitates a two-stage paradigm: query-agnostic memory construction followed by retrieval-based inference. Prior work invests in complex memory construction to pre-model high-level relations in videos, despite not knowing the downstream query at build time. We instead prioritize high-recall retrievability during memory building, and defer query-specific, high-level relation composition to inference time. To this end, we propose MERIT(Multi-key Episodic Retrieval with Inference-time Temporal expansion), a simple yet effective agentic framework for ultra-long video understanding. First, we formulate an episodic multi-key representation that enables precise retrieval of fine-grained memories through a simple key-matching mechanism. Second, we introduce a neighbor filtering mechanism to capture broader semantic context without the massive computational overhead of global memory construction. This is achieved by expanding the temporal scope exclusively around the retrieved segments at inference time. By leveraging simple key-matching with this on-demand temporal expansion, MERIT achieves state-of-the-art performance across three long-video benchmarks: EgoLifeQA, LVBench, and Video-MME (Long).

CommentsAccepted to ECCV 2026 (Oral). Project Page: https://choi-yeeun.github.io/MERIT/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑