arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13250cs.CVcs.AI

多模态大语言模型无关的即插即用关键帧选择方法用于长视频理解的评估

Evaluation of MLLM-Agnostic Plug-and-Play Keyframe Selection Methods for Long Video Understanding

Dilip Sarkar, Md. Safayet Islam, Liang Liang

首次发表
浏览论文内容

中文总结 AI 辅助

本文评估了五种无需训练的即插即用关键帧选择方法,在三个MLLM和三个长视频理解基准上,发现QAaF性能最佳,FOCUS次之。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)由于视觉令牌和计算预算的限制,无法处理长视频的每一帧。已提出三种主要方法来增强其长视频理解能力:(i)在大型视频语料库上重新训练MLLM和/或扩展其输入长度;(ii)为特定MLLM训练一个适配器,该适配器将整个视频和查询作为输入,并选择最相关的视频帧;(iii)开发一种无需训练、即插即用(PaP)且与MLLM无关的适配器。我们将第三种方法称为PaP关键帧选择。PaP方法可能仅使用候选视频帧而不考虑查询,也可能同时使用候选视频帧和查询。第一种方法成本过高。第二种方法需要大量的训练时间和计算资源,但由于适配器的可训练参数远少于整个MLLM,因此对许多人来说是可及的。第三种方法计算成本最低,因此广泛可及。据我们所知,在过去一年中仅报道了五种PaP方法。所有这些方法都在一个或多个视频问答基准上进行了评估,并展示了在长视频理解方面的改进。然而,这些方法是在不同的基准上使用不同的MLLM进行评估的。我们使用三种MLLM在三个长视频理解基准上对这五种方法进行了全面评估。我们的结果表明,QAaF在15个总体评估设置中的13个中取得了最佳性能,而FOCUS总体排名第二。这些结果为比较MLLM的无训练关键帧选择方法提供了共同的实验参考。

英文摘要

Multimodal large language models (MLLMs) cannot process every frame of a long video because of limitations in visual-token and computational budgets. Three primary approaches have been proposed to enhance their long-video understanding capabilities: (i) Retraining an MLLM on a large video corpus and/or extending its input length; (ii) Training an adapter for a specific MLLM that takes the entire video and the query as input and selects the most relevant video frames; and (iii) Developing a training-free, plug-and-play (PaP) adapter that is MLLM-agnostic. We refer to the third approach as PaP keyframe selection. A PaP method may use only candidate video frames without considering the query, or it may use both candidate video frames and the query. The first approach is prohibitively expensive. The second approach requires substantial training time and computational resources, but it is accessible to many because an adapter contains significantly fewer trainable parameters than an entire MLLM. The third approach has the lowest computational cost and is therefore broadly accessible. To the best of our knowledge, only five PaP methods have been reported within the past year. All of these methods have been evaluated on one or more video question-answering benchmarks and have demonstrated improvements in long-video understanding. However, the methods were evaluated on different benchmarks using different MLLMs. We present a comprehensive evaluation of these five methods using three MLLMs across three long-video understanding benchmarks. Our results show that QAaF achieves the best performance in 13 of the 15 aggregate evaluation settings, while FOCUS ranks second overall. These results provide a common experimental reference for comparing training-free keyframe-selection methods for MLLMs.

发表机构

  • University of Miami(迈阿密大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑