发表机构
Northwestern Polytechnical University(西北工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EC-RAG提出无需训练的事件链检索增强生成框架,通过构建显式事件链并融合多模态信号,在长视频理解中实现优于帧级检索的定位与推理性能。
AI 中文摘要
当前的大型视频-语言模型(LVLMs)在处理长视频时仍面临挑战,主要因为帧通常被独立处理,难以捕捉跨事件的时间依赖性。尽管检索增强方法已被引入以提供额外上下文,但大多数方法在帧或片段级别操作,这限制了它们对事件如何随时间演变及相互关联的建模能力。在本文中,我们提出事件链检索增强生成(EC-RAG),一种无需训练的框架,在问题回答之前将视频内容组织成显式的事件链。EC-RAG不是检索孤立的帧或文本片段,而是首先将视频划分为语义连贯的片段,使用多模态信号表示每个片段,然后将它们链接成保持时间顺序并捕捉事件间关系的结构化链。给定查询时,系统识别此链中的相关事件,并从关联模态中收集支持证据。我们的方法提供了几个实际优势:(i)事件级抽象更好地反映视频内容的自然结构,与帧级检索相比,实现更可靠的定位;(ii)结构化多模态融合在事件级别聚合语音、文本和视觉线索,使互补信息在推理过程中得到更有效的利用;(iii)与现有LVLM骨干的即插即用兼容性,无需额外训练或依赖专有模型。在Video-MME、MLVU和LongVideoBench上的实验表明,这种以事件为中心的设计始终优于帧级检索基线,突显了建模时间结构对长视频理解的重要性。
英文摘要
Current large video-language models (LVLMs) still face challenges when dealing with long videos, mainly because frames are often processed independently, making it difficult to capture temporal dependencies across events. Although retrieval-augmented approaches have been introduced to provide additional context, most of them operate at the frame or snippet level, which limits their ability to model how events evolve over time and relate to each other. In this paper, we propose Event Chain Retrieval-Augmented Generation (EC-RAG), a training-free framework that organizes video content into an explicit event chain before question answering. Instead of retrieving isolated frames or text segments, EC-RAG first partitions the video into semantically coherent segments, represents each segment using multi-modal signals, and then links them into a structured chain that preserves temporal order and captures inter-event relationships. Given a query, the system identifies relevant events within this chain and gathers supporting evidence from the associated modalities. Our approach offers several practical advantages: (i) event-level abstraction that better reflects how video content is naturally structured, enabling more reliable localization compared to frame-level retrieval; (ii) structured multi-modal fusion that aggregates speech, text, and visual cues at the event level, allowing complementary information to be more effectively utilized during reasoning; and (iii) plug-and-play compatibility with existing LVLM backbones, requiring no additional training or reliance on proprietary models. Experiments on Video-MME, MLVU, and LongVideoBench show that this event-centric design consistently outperforms frame-level retrieval baselines, highlighting the importance of modeling temporal structure for long-video understanding.
Comments12 pages, 7 figures, 7 tables, including supplementary material