arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

有证据支持的视频问答

Evidence-Backed Video Question Answering

Shijie Wang, Honglu Zhou, Ziyang Wang, Ran Xu, Caiming Xiong, Silvio Savarese, Chen Sun, Juan Carlos Niebles

arXiv 2607.11862首次发表:更新:

发表机构

Salesforce, Palo Alto, CA, USA; Brown University, Providence, RI, USA(Salesforce公司; 布朗大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视频问答中模型缺乏可视化依据的问题,提出E-VQA任务及ST-Evidence基准,开发数据集ST-Evidence-Instruct,通过微调提高模型表现,为可解释的视频理解建立基线。

AI 中文摘要

当前的视频大语言模型在问答方面表现出色,但大多像黑匣子一样运作,提供无可视化依据的文本答案。现有的可解释性方法依赖文本理由或稀疏边界框,难以捕捉复杂的视频动态。我们提出有证据支持的视频问答(E-VQA),要求模型联合输出语义答案和精确的时空证据。为此引入ST-Evidence基准,评估显示问答准确性和真正视觉感知之间存在关键解耦。我们开发数据集ST-Evidence-Instruct,在该数据上微调可提高模型表现,为可解释的视频理解建立了强大基线。

英文摘要

Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse bounding boxes, which struggle to capture complex video dynamics such as occlusions and non-rigid deformations. We propose Evidence-Backed Video Question Answering (E-VQA), a novel task requiring models to jointly output a semantic answer and precise spatio-temporal evidence: temporal segments and dense, tracked object segmentation masklets. To support this, we introduce ST-Evidence, the first human-verified benchmark for both discriminative and generative pixel-level grounding. Evaluations of state-of-the-art models reveal a critical decoupling between QA accuracy and true visual perception that scaling alone fails to bridge. To address this, we develop scalable, automated generation pipelines to create ST-Evidence-Instruct, a 160k-scale dataset bridging high-level reasoning with fine-grained grounding. Fine-tuning grounded Video LLMs on this data yields substantial gains over the corresponding size-matched UniPixel baselines (e.g., +27.2 t-mean and +13.8 J&F on a 7B model), establishing a robust baseline for explainable, evidence-backed video understanding. Code and data are available at https://github.com/SalesforceAIResearch/EVQA.

Journal refECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑