查询对齐的视频帧选择用于长视频理解
Query-aligned video frame selection for long video understanding
浏览论文内容
中文总结 AI 辅助
提出一种利用答案选项作为推理线索的查询对齐帧选择方法,通过子采样和余弦相似度评分,在固定令牌预算下提升长视频理解中多项选择题的视觉证据相关性。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)通过将文本、图像和视频转换为令牌序列,然后由骨干语言模型处理这些序列来处理多模态输入。尽管MLLMs在理解单张图像内容方面取得了优异性能,但视频理解仍然困难得多,因为视频包含大量视频帧。MLLMs通常只处理这些帧的子集,通常范围从8到64帧。MLLMs通常均匀采样帧,而不考虑它们与所回答问题的相关性。为了解决这一限制,最近提出了几种无需训练、模型无关的、选择与问题相关帧的方法。在这项工作中,我们介绍了一种专门为多项选择题设计的帧选择方法。我们通过附加从答案选项中提取的语义线索来扩展查询文本,并采用直接的查询-帧对齐评分机制。据我们所知,我们的方法是第一个直接利用答案选项作为推理时线索来选择与回答问题相关帧的方法。该方法首先通过以固定速率对视频帧进行子采样来构建一个紧凑的候选池。然后根据所有问答对的最大余弦相似度对帧进行评分,以识别给定查询最相关的帧。这种方法在保持固定令牌预算的同时,提高了提供给下游MLLM的视觉证据的相关性。我们在MLVU、Video-MME和LongVideoBench基准上使用三个MLLM:LLaVA-Mini、Qwen2-VL和LLaVA-Video评估了我们帧选择方法的有效性。实验结果表明,在相同帧预算下,答案感知的帧选择通常优于均匀采样和现有的无需训练的帧选择方法。
英文摘要
Multimodal large language models (MLLMs) process multimodal inputs by converting text, images, and videos into token sequences that are subsequently processed by a backbone language model. While MLLMs have achieved excellent performance in understanding the content of individual images, video understanding remains significantly more difficult because videos contain large number of video frames. MLLMs typically process only a subset of these frames, usually ranging from 8 to 64. MLLMs usually sample frames uniformly, regardless of their relevance to the question being answered. To address this limitation, several training-free, model-agnostic methods for selecting question-relevant frames have recently been proposed. In this work, we introduce a frame-selection method designed specifically for multiple-choice questions. We extend the query text by appending semantic cues derived from the answer choices and employ a direct query-frame alignment scoring mechanism. To the best of our knowledge, our method is the first to directly utilize answer choices as inference-time cues for selecting frames relevant to answering a question. The method first constructs a compact candidate pool by subsampling video frames at a fixed rate. The frames are then scored according to their maximum cosine similarity across all question-answer pairs to identify the most relevant frames for a given query. This approach preserves a fixed token budget while improving the relevance of the visual evidence provided to the downstream MLLM. We evaluate the effectiveness of our frame-selection method on the MLVU, Video-MME, and LongVideoBench benchmarks using three MLLMs: LLaVA-Mini, Qwen2-VL, and LLaVA-Video. Experimental results demonstrate that answer-aware frame selection generally outperforms uniform sampling and existing training-free frame-selection methods under the same frame budget.
发表机构
- University of Miami(迈阿密大学)
机构由 AI 辅助整理,请以论文原文为准。