arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23330cs.CV

IntentQA:基于认知上下文推理的视频意图问答

IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning

Jiapeng Li, Ping Wei, Wenjuan Han, Song-Chun Zhu, Lifeng Fan

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出视频意图问答任务 IntentQA,构建相关大规模数据集,提出 X-CaVIR 框架并引入对比性能下降指标,实验验证其有效性、优越性与稳定性。

中文摘要 AI 辅助

视频理解要求智能体超越单纯的视觉事实识别,理解人类行为背后的意图(常被称为社会智能的“暗物质”)。为弥合视觉观察与意图推理之间的差距,本文提出了一项新任务 IntentQA,并为此贡献了一个大规模的视频问答(VideoQA)数据集。考虑到标准指标可能因数据集偏差而高估模型能力,本文超越了简单的准确率,对模型的鲁棒性进行严格评估。本文通过大语言模型(LLMs)生成五个不同的对比集来扩充基准,并引入“对比性能下降”指标。本文提出了 X-CaVIR(可解释的上下文感知视频意图推理)框架,该框架利用三类“认知上下文”来增强视频分析:i)通过跨模态视频查询语言(VQL)模块实现的情境上下文;ii)通过对比学习模块实现的对比上下文;iii)通过常识推理模块实现的常识上下文。至关重要的是,为克服传统黑箱模型的不透明性,本文通过采用透明流水线优化了 X-CaVIR 中 LLMs 的集成,该流水线将视频字幕与视觉问答(VQA)模型输出协同结合。这种方法不仅通过有效利用丰富的常识知识提升了性能,还使推理过程明确可解释。大量实验证明了本文各组件的有效性、X-CaVIR 相较于现有最优基准的优越性,以及其在对比集扰动下的稳定性。

英文摘要

Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions (often termed the "dark matter" of social intelligence). To bridge the gap between visual observation and intent reasoning, we introduce a novel task, IntentQA, and contribute a large-scale VideoQA dataset specifically tailored for this purpose. However, recognizing that standard metrics may overestimate capabilities due to dataset biases, we go beyond simple accuracy to rigorously evaluate model robustness. We augment the benchmark by generating five distinct contrast sets via Large Language Models (LLMs) and introducing a "Contrast Performance Decline" metric. We propose the X-CaVIR (eXplainable Context-aware Video Intent Reasoning) framework, which leverages three types of "Cognitive Context" to enhance video analysis: i) Situational Context via a cross-modal Video Query Language (VQL) module, ii) Contrastive Context via a Contrastive Learning module, and iii) Commonsense Context via a Commonsense Reasoning module. Crucially, to overcome the opacity of traditional black-box models, we refine the integration of LLMs within X-CaVIR by employing a transparent pipeline that synergizes video captions with VQA model outputs. This approach not only improves performance by effectively utilizing rich commonsense knowledge but also renders the reasoning process explicitly interpretable. Extensive experiments demonstrate the effectiveness of our components, the superiority of X-CaVIR over state-of-the-art baselines, and its stability against perturbations on the contrast sets.

发表机构

  • Beijing Institute for General Artificial Intelligence (BIGAI)(北京通用人工智能研究院)
  • Beijing Jiaotong University(北京交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑