arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15708cs.CV

你所问即你所定位:连接问题意图与时序证据以实现定位视频问答

What You Ask is What You Ground: Bridging Question Intent to Temporal Evidence for Grounded VideoQA

  • KAIST(韩国科学技术院)
  • Ewha Womans University(梨花女子大学)

机构由 AI 辅助整理,请以论文原文为准。

Jinhwan Seo, Kyubeom Han, Jumin Lee, Junhyug Noh, Sung-eui Yoon

AI总结:

针对定位视频问答中问题不变定位的失效模式,提出GroundFormer模型,通过可学习通信令牌、因子化MIL交叉注意力等技术,在NExT-GQA和STAR数据集上实现最优性能,提升时序定位的问题区分性。

AI中文摘要:

我们研究了定位视频问答(Grounded Video Question Answering)中一种关键却被忽视的失效模式:问题不变定位,即模型对同一视频的不同问题预测出几乎相同的时序片段。我们将此行为归因于现有常见设计的两个结构局限:(i)模态隔离,即视频表征在接收问题语义前就已固定;(ii)定位模块内的问题注入较弱。为解决这些问题,我们提出了GroundFormer,该模型通过可学习的通信令牌介导定向视觉-语言交互,在定位前使视频特征受问题意图调控。在问题调控的特征之上,因子化多实例学习(MIL)交叉注意力在候选级监督下将答案选择与时序证据结合,同时高斯平滑将峰值注意力转换为时间连贯的片段。我们进一步引入分层多模态对比损失,在双次训练流程中对齐视频、问题和答案嵌入。GroundFormer在NExT-GQA和STAR数据集上实现了定位视频问答的最优性能,大幅提升了具问题区分性的时序定位效果。

英文摘要:

We study a critical yet overlooked failure mode in Grounded Video Question Answering: question-invariant grounding, where models predict nearly identical temporal segments for different questions about the same video. We trace this behavior to two structural limitations in prior common designs: (i) modality isolation that fixes video representations before they receive question semantics, and (ii) weak question injection inside the grounding module. To address this, we propose GroundFormer, which conditions video features on question intent before localization via learnable communication tokens that mediate directed visuo-lingual interaction. On top of the question-conditioned features, a factorized MIL cross-attention couples answer selection with temporal evidence under candidate-level supervision, while Gaussian smoothing converts peaked attention into temporally coherent segments. We further introduce a hierarchical multi-modal contrastive loss that aligns video, question, and answer embeddings across a two-pass training pipeline. GroundFormer achieves state-of-the-art grounded VideoQA performance on NExT-GQA and STAR, substantially improving question-discriminative temporal grounding.

补充信息

↑