发表机构
Zhejiang University; South China University of Technology; Lenovo Group Limited(浙江大学; 华南理工大学; 联想集团有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出以事件为中心的多模态记忆框架EM^2Mem,将异质证据绑定到事件锚点,在三个长视频问答基准上提升了准确率、召回率,降低了延迟和推理token用量。
AI 中文摘要
多模态记忆为长视频问答提供了可扩展的接口,但现有方法通常检索字幕、帧、文字记录、摘要或图事实作为孤立片段。尽管可检索,但这类片段不具备生成就绪性:语言模型必须在推理时重构跨模态和时间对齐,此时上下文有限且归因困难。我们提出EM^2Mem,一种以事件为中心的多模态记忆框架,该框架在记忆构建期间将异质证据绑定到事件锚点。每个以事件索引的记忆单元对齐多模态记录、时间上下文、图关联、语义事实和来源,支持对基于接地多模态事件的紧凑证据读出,而非针对特定模态的片段。在三个长视频问答基准上,EM^2Mem较最强记忆基线的平均准确率分别提升2.0、2.4和3.7个百分点,严格事件级Top-5证据召回率提升7.0个百分点,每个查询的延迟降低4.67倍,总推理 token 减少63.66%(代码将集成至该httpsURL)。
英文摘要
Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as isolated fragments. Although searchable, such fragments are not generation-ready: language models must reconstruct cross-modal and temporal alignments at inference time, when context is limited and attribution is difficult. We propose EM^2Mem, an event-centric multimodal memory framework that binds heterogeneous evidence to event anchors during memory construction. Each event-indexed memory cell aligns multimodal records, temporal context, graph-linked relations, semantic facts, and provenance, enabling compact evidence readout over grounded multimodal events rather than modality-specific fragments. Across three long-video QA benchmarks, EM^2Mem improves average accuracy over the strongest memory baseline by 2.0, 2.4, and 3.7 points, improves strict event-level Top-5 evidence recall by 7.0 points, and reduces per-query latency by 4.67 times and total inference tokens by 63.66% (The code will be integrated into https://github.com/zjunlp/LightMem).
CommentsAccepted by EMNLP 2026 findings