发表机构
Georgia Institute of Technology; Center for Signal and Information Processing (CSIP); School of Electrical and Computer Engineering(佐治亚理工学院; 信号与信息处理中心 (CSIP); 电气与计算机工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究长视频理解中推理时最大化查询相关信息问题,提出FORGE模型无关方法,通过在预训练嵌入空间引入查询条件几何统一相关性和多样性,实验证明该方法在关键帧选择和问答上有显著提升。
AI 中文摘要
多模态大语言模型使长视频理解成为可能,但随着视频序列长度增加,相关内容密度急剧下降,更多无关内容会降低模型准确性。本文解决在推理时选择的帧子集中最大化查询相关信息的问题。FORGE是一种模型无关方法,在预训练多模态嵌入空间上引入查询条件几何,将相关性和多样性统一为单个目标。实验表明,FORGE在关键帧选择得分上比最强无训练基线提高11.0 - 15.3分,关键帧召回率提高一倍。在问答中,跨8个开源MLLM的评估设置中准确性提高,比均匀采样高8.7分,比最强基线高5.2分。研究表明使嵌入空间与查询高维结构对齐是推理时视频理解的有前景方向。
英文摘要
Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible. However, the density of relevant content decreases sharply as video sequence length increases, and exposing the model to more irrelevant content measurably reduces its accuracy. In this paper, we address the problem of maximizing query-relevant information in a frame subset selected at inference time, without training. FORGE (Frame Orthogonality in Relevance Geometry) is a model-agnostic method that induces a query-conditioned geometry on a pretrained multimodal embedding space, unifying relevance and diversity into a single objective. In this space, frames that cover independent query-relevant directions are far apart, and selecting the subset of maximum information captures diverse query-relevant content within the budget. Experiments on Video-MME and LongVideoBench at budgets of 16, 32, and 64 frames show that FORGE improves the unified keyframe selection score by 11.0-15.3 points over the strongest training-free baseline and up to doubles keyframe recall (0.415 vs. 0.204 at K=64 on Video-MME). The gains extend to question answering, where accuracy improves in every evaluated setting across eight open-source MLLMs spanning 4B to 32B parameters, by up to 8.7 points over uniform sampling and 5.2 points over the strongest baseline. Our findings suggest that aligning the embedding space with the query's high-dimensional structure is a promising direction for inference-time video understanding.
CommentsUnder Review