arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向视频理解的覆盖驱动型自适应关键帧选择

Coverage-Driven Adaptive Keyframe Selection for Video Understanding

Junyang Zhang, Puhan Luo, Chen Tang, Yuxi Shi, Xiang-Yang Li

arXiv 2608.00714首次发表:更新:

AI 中文总结

本文针对视频理解中LVLM处理大量帧计算开销大的问题,提出无需训练的CSES关键帧选择器,通过覆盖驱动的自适应策略减少需评分和选择的帧数量,同时保持准确率并提升选择速度。

AI 中文摘要

近期大型视觉语言模型(LVLMs)的进展已实现长视频理解与分析,但处理视频中大量帧会产生巨大计算开销。现有方法通过推理前对帧-查询相关性评分并据此选择关键帧,以降低LVLM推理成本。然而,相关帧的分布随查询变化,这些方法常需对数百或数千帧评分。为解决此局限,本文提出CSES,一种无需训练的语义关键帧选择器,可自适应确定需评分的帧数量与需选择的关键帧数量。CSES估计帧-查询相关性分布的显著性,以指导主动获取并调整各输入的时间覆盖范围,随后将关键帧选择建模为覆盖问题,联合考虑语义相关性、时间冗余与视觉冗余。主动获取与关键帧选择基于覆盖饱和终止。选择目标是单调子模函数,可通过标准近似保证实现贪心优化。在两个基准上对四种LVLMs开展的实验表明,本文方法在保持准确率的同时,需评分的帧比现有基线少4-13倍,选择的输入关键帧少18.4%-20.5%,帧选择速度较基线进一步提升3.1-5.4倍。

英文摘要

Recent advances in large vision-language models (LVLMs) have enabled long-video understanding and analysis. However, processing the large number of frames in a video incurs substantial computational overhead. Existing methods reduce LVLM inference costs by scoring frame-query relevance before inference and selecting keyframes accordingly. Nevertheless, the distribution of relevant frames varies across queries, and these methods often need to score hundreds or thousands of frames. To address this limitation, we propose CSES, a training-free semantic keyframe selector that adaptively determines the numbers of frames to score and keyframes to select. CSES estimates the prominence of the frame-query relevance profile to guide active acquisition and adapt the temporal coverage of each input. It then formulates keyframe selection as a coverage problem that jointly accounts for semantic relevance, temporal redundancy, and visual redundancy. Active acquisition and keyframe selection terminate based on coverage saturation. The selection objective is monotone and submodular, enabling greedy optimization with a standard approximation guarantee. Experiments with four LVLMs on two benchmarks show that our method preserves accuracy while scoring $4$-$13\times$ fewer frames and selecting $18.4\%$-$20.5\%$ fewer input keyframes than existing baselines. CSES further achieves a $3.1$-$5.4\times$ speedup in frame selection over baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑