arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探究时空视频定位中的查询不敏感行为

Investigating Query-Insensitive Behavior in Spatio-Temporal Video Grounding

Eryk Kołodziejczyk, Alberto Presta, Karol Szurkowski, Michal Byra

arXiv 2610.06018首次发表:更新:

发表机构

Samsung AI Center(三星人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究挑战STVG模型假设查询与视频相关的传统,发现现有模型在无关或缺失查询时仍能预测,分析数据集规律并提出负感知评估协议与架构以评估查询相关性。

AI 中文摘要

时空视频定位(STVG)旨在定位自然语言查询所描述的对象或事件在空间和时间上的位置。现有STVG模型通常在假设每个查询与输入视频相关的条件下进行训练和评估。在本工作中,我们通过研究最先进的STVG模型在无关查询和缺失文本输入下的行为来挑战这一假设。我们的实验表明,即使查询与视频无关或完全移除,当前模型仍能产生合理的时空预测。我们进一步分析HCSTVG-v2和VidSTG,以识别可能鼓励这种查询不敏感行为的数据集规律。我们的研究凸显了STVG模型一个未被充分探索的局限性,并促使采用明确评估查询相关性的负感知评估协议和架构。

英文摘要

Spatio-temporal video grounding (STVG) aims to localize objects or events described by natural language queries in both space and time. Existing STVG models are typically trained and evaluated under the assumption that each query is relevant to the input video. In this work, we challenge this assumption by studying the behavior of state-of-the-art STVG models under irrelevant queries and missing textual input. Our experiments show that current models can still produce plausible spatio-temporal predictions even when the query is unrelated to the video or removed entirely. We further analyze HCSTVG-v2 and VidSTG to identify dataset regularities that may encourage such query-insensitive behavior. Our study highlights an underexplored limitation of STVG models and motivates negative-aware evaluation protocols and architectures that explicitly assess query relevance.

CommentsAccepted on EMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑