arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FORTE:面向长视频问答的自适应评分与精确关键帧选择

FORTE: Adaptive Scoring and Exact Keyframe Selection for Long-Video Question Answering

Haifeng Huang, Biyin Xu, Chunsheng Xin, Yang Li

arXiv 2610.00573首次发表:更新:

发表机构

Iowa State University(爱荷华州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FORTE提出一种无需训练的自适应评分与精确关键帧选择框架,通过高斯过程预测和全局优化,在有限预算下高效提升长视频问答的准确率。

AI 中文摘要

查询感知的关键帧选择使多模态大语言模型(MLLMs)能够仅使用一小部分与问题相关的帧来处理长视频。然而,现有的基于评分的方法通常在固定的、均匀采样的候选池内进行搜索,从而使得该池之外的证据永远无法被选中。在有限的关联性评分预算下,关键挑战在于如何自适应地将评估分配给有前景的帧,同时继续探索代表性不足的时间区域。我们提出了FORTE,一个无需训练的框架,通过两个阶段解决这一挑战:自适应关联性评分和全局关键帧优化。从稀疏、均匀分布的观测开始,我们高效的基于高斯过程的关联性预测器估计未评分帧的关联性,利用时间局部性和近似带状核结构,将固定带宽下帧数核心计算从立方时间降至线性时间。随后,评分阶段通过平衡预测关联性与时间覆盖度来选择接下来要评分的帧,优先考虑有前景的区域,同时探索视频中代表性不足的部分。优化阶段通过最大化一个联合捕获测量关联性和时间覆盖度的目标来选择最终关键帧。我们推导出一种精确算法,利用对数覆盖结构,在固定最终帧预算下,以与池大小成线性时间识别评分候选池中的最优子集。在四个长视频问答基准上的实验表明,在每种测试的评分预算下,FORTE在比较的选择器中实现了最高的观测平均准确率。进一步的评估证明了其在不同关联性评分器和下游MLLMs中的一致有效性。

英文摘要

Query-aware keyframe selection enables multimodal large language models (MLLMs) to process long videos using only a small set of question-relevant frames. Existing score-based methods, however, typically search within a fixed, uniformly sampled candidate pool, preventing evidence outside this pool from ever being selected. Given a limited relevance-scoring budget, the key challenge is to allocate evaluations adaptively to promising frames while continuing to explore underrepresented temporal regions. We introduce FORTE, a training-free framework that addresses this challenge through two stages: adaptive relevance scoring and global keyframe optimization. Starting from sparse, uniformly distributed observations, our efficient Gaussian-process relevance predictor estimates relevance for unscored frames, exploiting temporal locality and the approximately banded kernel structure to reduce the core computation from cubic to linear time in the number of frames for fixed bandwidth. The scoring stage then selects which frames to score next by balancing predicted relevance with temporal coverage, prioritizing promising regions while also exploring less-represented parts of the video. The optimization stage selects the final keyframes by maximizing an objective that jointly captures measured relevance and temporal coverage. We derive an exact algorithm that leverages the logarithmic coverage structure to identify the optimal subset of the scored candidate pool in time linear in the pool size, for a fixed final-frame budget. Experiments on four long-video question-answering benchmarks show that FORTE achieves the highest observed mean accuracy among the compared selectors under every tested scoring budget. Further evaluations demonstrate its consistent effectiveness across different relevance scorers and downstream MLLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑