arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31873cs.CVcs.AIcs.CL

CueKFS:面向长视频理解的智能体提示驱动关键帧选择

CueKFS: Agentic Cue-Driven Keyframe Selection for Long Video Understanding

Weitai Kang, Hanieh Deilamsalehy, Yumo Xu, Dewang Sultania, Serdar Cellat, Yan Yan

首次发表
浏览论文内容

中文总结 AI 辅助

CueKFS提出无需训练的关键帧选择方法,通过动态生成视觉提示并智能体式修订,实现长视频问答的最优性能,增益最高达4.54%。

中文摘要 AI 辅助

关键帧选择(KFS)长期以来为浏览和检索生成紧凑的视频摘要,并为缩略图提供代表性帧。最近,在给定问题的条件下,KFS通过选择与问题更相关的帧,为长视频问答提供了一种替代均匀采样的方案。大多数方法根据帧与问题的相似度进行排序。然而,当问题组合了单个帧无法同时展示的主体或时刻,或需要措辞中未包含的隐含信息时,相关帧的得分可能很低。其他方法尝试将问题分解为子查询,但由于静态的初始上下文,分解往往不准确。因此,我们提出了CueKFS,一种无需训练的方法,将问题-帧匹配重新表述为将帧与一组动态生成的视觉提示进行比较。从一组初始显著帧出发,我们将问题分解为提示。每个提示并发地探测视频以定位其自身的证据。然后,一个推理型视觉语言模型(VLM)智能体地根据证据修订提示集,以重新探索视频。CueKFS随后在存活的提示之间分配预算。在三个基准测试中,CueKFS在所有27个具有现有先前结果的评估设置中取得了最先进的结果,与之前的基线相比,预算平均增益最高达+4.54%,且中位数仅需两次VLM调用。我们进一步提供了CueKFS的详细行为分析,表明智能体提示细化驱动了对视频的主动重新探索,相对于初始上下文产生了高达92%的相对相似度增益。

英文摘要

Keyframe selection (KFS) has long produced compact video summaries for browsing and retrieval, and representative frames for thumbnails. More recently, when conditioned on a question, KFS provides an alternative to uniform sampling for long-video question answering by selecting frames that are more relevant to the question. Most methods rank frames by similarity to the question. Yet a relevant frame may score poorly when the question combines subjects or moments that no single frame shows, or requires implicit information absent from its wording. Other methods try to break down the question into subqueries, but suffer from inaccurate decomposition due to their static initial context. Therefore, we propose CueKFS, a training-free method that reformulates question--frame matching as comparing frames against a set of dynamically generated visual cues. From an initial set of salient frames, we decompose the question into cues. Each cue concurrently probes the video to navigate to its own evidence. A reasoning VLM then agentically revises the cue set against its evidence to re-explore the video. CueKFS then allocates the budget across the surviving cues. Across three benchmarks, CueKFS establishes state-of-the-art results in all 27 evaluated settings with available prior results, achieving budget-averaged gains of up to +4.54% over the previous baseline and a median of only two VLM calls. We further provide a detailed behavioral analysis of CueKFS, showing that agentic cue refinement drives active re-exploration of the video, yielding relative similarity gains of up to 92% over the initial context.

补充信息

↑