发表机构
The University of Sydney; University of Amsterdam(悉尼大学; 阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频时间定位中动作与实体未显式对齐的问题,提出IAE-VTG,通过细粒度解耦交互模块和交互敏感分配,在表示与训练分配层面建模交互,提升复杂事件定位性能。
AI 中文摘要
视频时间定位(VTG)旨在定位与自然语言查询相匹配的视频片段。许多查询描述的是特定实体执行的动作。现有方法通常将查询整体编码或使用通用的视频-文本交互,而没有显式检查动作和实体是否共同出现。因此,它们可能选择同时包含这两个概念但并非查询所描述事件的片段。我们提出了交互对齐的动作-实体视频时间定位(IAE-VTG),该方法在表示层面和训练分配层面均对此关系进行建模。首先,细粒度解耦交互模块(FDIM)将查询信息分离为动作相关和实体相关部分,并将其与互补的运动和外观特征对齐。随后,该模块结合词元级交互来构建捕获动作与实体之间关系的表示。其次,交互敏感分配(ISA)将这种交互证据添加到二分图匹配中,使得训练目标的选择同时基于时间重叠和语义兼容性。这减少了来自时间上合理但语义上不正确的提案的监督。在QVHighlights、Charades-STA和TACoS上的实验表明,IAE-VTG持续改进强基线,并在标准定位指标上达到具有竞争力或最先进的性能。进一步的分析表明,当相似动作或实体在多个时间出现时,该方法尤其有效,并为复杂事件产生更可靠的分配。
英文摘要
Video Temporal Grounding (VTG) localizes the video segment that matches a natural-language query. Many queries describe an action performed by a particular entity. Existing methods often encode the query as a whole or use general video-text interactions, without explicitly checking whether the action and entity occur together. They may therefore select a segment that contains both concepts but not the event described by the query. We propose Interaction Aligned Action-Entity Video Temporal Grounding (IAE-VTG), which models this rela?tionship at both the representation and training assignment levels. First, the Fine-grained Disentangled Interaction Module (FDIM) separates action and entity related query information and aligns it with complementary motion and appearance features. It then combines token-level interactions to build representations that capture the relationship between the action and entity. Second, Interaction-Sensitive Assignment (ISA) adds this interaction evidence to bipartite matching, so training targets are selected using both temporal overlap and semantic compatibility. This reduces supervision from temporally plausible but semantically incorrect proposals. Experiments on QVHighlights, Charades?STA, and TACoS show that IAE-VTG consistently improves strong baselines and achieves competitive or state-of-the-art performance on standard grounding metrics. Additional analyses show that the method is especially effective when similar actions or entities appear at multiple times and produces more reliable assignments for complex events.