发表机构
Department of Computer Science and Engineering, HKUST(香港科技大学计算机科学与工程学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长视频问答的证据稀疏、选项区分困难问题,提出因子引导的PACE框架,经实验在多基准数据集上优于现有方法,实现性能提升。
AI 中文摘要
尽管大型视觉语言模型(LVLM)发展迅速,长视频问答仍然存在挑战:相关证据稀疏,与问题相关的上下文往往无法提供线索以区分正确答案与看似合理的替代选项。对MMR-V手动注释子集的诊断分析显示,现有智能体系统相比直接视觉语言模型(VLM)推理在线索检索上有显著提升,但未实现答案准确率的相应增长,这表明瓶颈在于区分选项的证据,而非仅主题相关性。我们提出PACE(关键证据的渐进获取,Progressive Acquisition of Critical Evidence),这是一种用于长视频证据获取的因子引导框架。PACE分为两个阶段:第一阶段在不观察候选答案的情况下,基于问题导出的因子对片段级描述进行索引;第二阶段利用候选答案推导对比线索,并查询索引进行验证。在采用开源Qwen3-VL骨干模型的MMR-V上,PACE达到42.6%的准确率,优于直接推理及包括Deep Video Discovery(DVD)在内的现有智能体基线。在同一诊断子集上,PACE恢复了66.9%的注释线索,提供了经验证据表明其性能提升与证据恢复的改进相关,而非仅依赖更强的答案侧先验。在LVBench、Video-MME、EgoSchema和LongVideoBench上,PACE相比DVD均实现一致提升,表明选项感知的证据获取能力可迁移至MMR-V之外的数据集。代码可在指定URL获取。
英文摘要
While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-relevant context often fails to provide cues that discriminate the correct answer from plausible alternatives. Diagnostic analysis on a manually annotated subset of MMR-V shows that prior agentic systems substantially improve cue retrieval over direct VLM inference yet fail to achieve a corresponding gain in answer accuracy, indicating that the bottleneck lies in option-discriminative evidence rather than topical relevance alone. We propose PACE (Progressive Acquisition of Critical Evidence), a factor-guided framework for long-video evidence acquisition. PACE proceeds in two stages: it first indexes clip-level descriptions guided by question-derived factors without observing the candidate answers; it then uses the candidate answers to derive contrastive cues and queries the index for verification. On MMR-V with the open-source Qwen3-VL backbone, PACE achieves 42.6% accuracy, outperforming direct inference and prior agentic baselines including Deep Video Discovery (DVD). On the same diagnostic subset, PACE recovers 66.9% of the annotated cues, providing empirical evidence that its gains are associated with improved evidence recovery rather than stronger answer-side priors alone. Consistent gains over DVD on LVBench, Video-MME, EgoSchema, and LongVideoBench suggest that option-aware evidence acquisition transfers beyond MMR-V. Code is available at https://github.com/HKUST-KnowComp/PACE.