arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MarKey:边际效用引导的贪心关键帧选择用于长视频理解

MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding

Hongchang Shi, Jinpeng Hu, Ao Wang, Wenzheng Zhou, Hui Ma, Feng Li, Zenglin Shi

arXiv 2609.15408首次发表:更新:

发表机构

Hefei University of Technology(合肥工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长视频理解中关键帧选择冗余和证据覆盖不全的问题,提出免训练框架MarKey,通过子集感知的贪心优化联合考虑查询相关性、边际覆盖增益和冗余,在六个基准上持续优于现有方法。

AI 中文摘要

长视频理解对多模态大语言模型(MLLMs)而言仍然具有挑战性,因为密集编码长帧序列的计算成本高昂,而在有限的视觉预算下进行均匀采样可能会遗漏稀疏但决定性的证据。近年来的免训练关键帧选择方法实现了更高效的推理,并取得了有前景的性能提升。然而,许多现有方法在很大程度上孤立地对帧进行评分,没有明确考虑每个候选帧如何与当前已选子集互补,可能导致冗余选择和证据覆盖不完整。为解决这一局限,我们提出MarKey,一个免训练框架,将关键帧选择表述为子集感知的贪心优化。在每次迭代中,MarKey使用一个易处理的代理函数对每个候选帧进行评分,该代理函数联合考虑查询相关性、边际覆盖增益和上下文相关冗余,并选择效用最高的帧。为使这种迭代子集感知评估高效,MarKey使用一组紧凑的代表性锚点来近似全视频覆盖,并使用一个有限的先前已选帧窗口来限制上下文相关比较。在六个基准上的实验涵盖整体视频理解、以人为中心的视频理解和开放式视频理解,结果表明MarKey持续优于现有方法。进一步分析显示,在不同MLLM骨干网络、模型规模和帧预算下均取得了稳健的性能提升。

英文摘要

Long-video understanding remains challenging for multimodal large language models (MLLMs) because densely encoding long frame sequences is computationally expensive, while uniform sampling under a limited visual budget can miss sparse yet decisive evidence. Recent training-free keyframe selection methods have enabled more efficient inference and yielded promising performance gains. However, many existing methods score frames largely in isolation without explicitly considering how each candidate complements the currently selected subset, potentially resulting in redundant selections and incomplete evidence coverage. To address this limitation, we propose MarKey, a training-free framework that formulates keyframe selection as subset-aware greedy optimization. At each iteration, MarKey scores each candidate using a tractable surrogate that jointly accounts for query relevance, marginal coverage gain, and context-dependent redundancy, and selects the frame with the highest utility. To make this iterative subset-aware evaluation efficient, MarKey uses a compact set of representative anchors to approximate full-video coverage and a bounded window of previously selected frames to limit context-dependent comparisons. Experiments on six benchmarks spanning holistic video understanding, human-centric video understanding, and open-ended video understanding demonstrate that MarKey consistently outperforms existing methods. Further analyses show robust gains across different MLLM backbones, model scales, and frame budgets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑