发表机构
National University of Singapore; Institute of High Performance Computing, Agency for Science, Technology and Research (A*STAR); Nanyang Technological University; University of Science and Technology of China; Jilin University(新加坡国立大学; 高性能计算研究所,科学、技术与研究机构(A*STAR); 南洋理工大学; 中国科学技术大学; 吉林大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对多模态大语言模型长视频理解受限问题,提出ReMem框架,通过双级记忆增强自适应,在查询和视频级别分别处理,能适应不同时间粒度,经实验验证在多个基准测试中实现高效零样本性能,提升模型长视频推理能力。
AI 中文摘要
虽然多模态大语言模型(MLLMs)在基本视频任务中表现出卓越的泛化能力,但受限的上下文窗口限制了它们对长视频的理解。为适应这一限制,模型通常采用关键帧选择。然而,均匀采样或静态查询引导选择往往忽略关键时间上下文,无法适应不同的查询时间粒度。本文提出了ReMem,一种用于免训练长视频问答(LongVideoQA)的时间粒度自适应关键帧选择框架。ReMem引入了双级记忆增强自适应。在查询级别,记忆驱动问题解析利用大语言模型的长期记忆来解码问题时间粒度并提取语义实体。在视频级别,协同双语义帧对齐利用内在结构记忆将帧与查询语义对齐,指导结构感知动态帧路由对事件进行聚类并优化分配采样预算。通过记忆机制明确保留时间信息,ReMem抑制冗余并使MLLMs能够进行强大的多粒度视频推理。使用三个MLLMs在四个流行的LongVideoQA基准上的评估证明了其高效的、当前最优的零样本性能;值得注意的是,配备ReMem的LLaVA - Video在LVBench上达到54.5%(+12.3%),在LongVideoBench上达到67.1%(+8.2%)。
英文摘要
While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding. To accommodate this constraint, models typically resort to keyframe selection. However, uniform sampling or static query-guided selection often overlooks critical temporal context, failing to adapt to the varying query temporal granularities. In this paper, we propose ReMem, a temporal granularity-adaptive keyframe selection framework for training-free LongVideoQA. ReMem introduces a dual-level memory-augmented adaptation. At the query level, Memory-Driven Question Parsing leverages LLM long-term memory to decode question temporal granularity and extract semantic entities. At the video level, Synergistic Dual-Semantic Frame Alignment exploits intrinsic structural memory to align frames with query semantics, guiding Structure-Aware Dynamic Frame Routing to cluster events and optimally distribute sampling budgets. By explicitly preserving temporal information with memory mechanisms, ReMem suppresses redundancy and empowers MLLMs to perform robust multi-granular video reasoning. Evaluations across four popular LongVideoQA benchmarks using three MLLMs demonstrate highly efficient, state-of-the-art zero-shot performance; notably, LLaVA-Video with ReMem reaches 54.5% (+12.3%) on LVBench and 67.1% (+8.2%) on LongVideoBench.
CommentsAccepted by ECCV 2026