AI 中文总结
本文针对长视频问答中视觉预算有限及证据对齐弱的问题,提出无需训练的GCR框架,通过锚定、覆盖、优化三步选择帧,在LongVideoBench等基准上较基线取得显著性能提升。
AI 中文摘要
长视频问答需要在有限的视觉令牌预算下,从包含数千帧的视频中识别出稀疏但关键的证据。现有方法要么单次选择与查询相关的帧,要么仅依赖带时间戳的文本作为检索指导,存在两个关键局限:一是所选帧往往集中在局部相关性峰值处,预算耗尽后遗漏的证据无法恢复;二是文本与视觉证据的对齐度较弱。本文提出GCR,这是一个无需训练的框架,将固定预算的帧选择转化为联合证据整理问题:锚定(Ground)将带时间戳的文本转换为时间事件,选择与查询相关的真实帧锚点,并将每个事件文本渲染到其时间对齐的帧上;覆盖(Cover)为已锚定的事件补充直接的视觉锚点以获取互补视觉证据,并应用全局最大边际相关性以保留多样化上下文;优化(Refine)重新访问遗漏的时间区域,用真实帧的 medoid(类中心)替换最弱的可修订上下文帧——仅当该类中心能提供更高的证据价值时才执行此操作。GCR保持固定数量的按时间顺序排列的帧,无需进行视觉语言模型(VLM)训练或架构修改。在LongVideoBench和Video-MME数据集上,使用三个7B规模的骨干模型,帧预算为8、32和64时的实验表明,GCR在长视频问答任务中取得了一致的性能提升。采用7B LLaVA-OV骨干模型和32帧设置时,GCR在两个基准上分别达到64.25%和62.15%的准确率,分别比最强的复现基线高出2.54和1.93个百分点。
英文摘要
Long-video question answering requires identifying sparse yet critical evidence from videos containing thousands of frames under a constrained visual-token budget. Existing methods either select query-aware frames in a single pass or rely on timestamped text solely as retrieval guidance, leading to two key limitations. First, selected frames tend to cluster around local relevance peaks, and once the budget is exhausted, omitted evidence cannot be recovered. Second, textual and visual evidence remain weakly aligned. We propose GCR, a training-free framework that casts fixed-budget frame selection as a joint evidence curation problem. Ground converts timestamped text into temporal events, selects query-relevant real frame anchors, and renders each event text onto its temporally aligned frame. Cover supplements grounded events with direct visual anchors for complementary visual evidence and applies global maximal marginal relevance to preserve diverse context. Refine revisits omitted temporal regions and replaces the weakest revisable context frame with a real-frame medoid---but only when the medoid offers greater evidence value. GCR maintains a fixed number of chronologically ordered frames and requires no VLM training or architectural modification. Experiments on LongVideoBench and Video-MME, across three 7B backbones and frame budgets of 8, 32, and 64, demonstrate consistent improvements in long-video QA. With the 7B LLaVA-OV backbone and 32 frames, GCR achieves 64.25% and 62.15% on the two benchmarks, outperforming the strongest reproduced baselines by 2.54 and 1.93 percentage points, respectively.