双任务互增强嵌入的联合视频段落检索与定位
Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding
- Kunming University of Science and Technology(昆明理工大学)
- Harbin Institute of Technology Shenzhen(哈尔滨工业大学深圳校区)
- Yunnan University(云南大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对视频段落定位依赖大规模标注时间标签及已知视频-段落对应关系的问题,提出DMR-JRG方法,通过检索与定位双任务互增强,减少模态差异并构建粗细粒度特征空间,大幅降低标注需求并实现精准匹配定位。
AI中文摘要:
视频段落定位(VPG)旨在从视频中精准定位与给定文本段落查询最相关的时刻。然而,现有方法通常依赖大规模标注的时间标签,并假设视频与段落的对应关系已知。这在实际应用中不切实际,因为构建时间标签需要大量人力成本,且对应关系往往未知。为解决此问题,我们提出双任务互增强嵌入的联合视频段落检索与定位方法(DMR-JRG)。该方法中检索与定位任务相互增强,而非独立处理。DMR-JRG主要包含检索分支和定位分支:检索分支通过视频间对比学习大致对齐段落与视频的全局特征,减少模态差异并构建粗粒度特征空间,摆脱段落与视频对应关系的依赖;此粗粒度特征空间进一步辅助定位分支提取细粒度上下文表示,定位分支通过探索视频片段与文本段落的局部、全局及时间维度一致性,实现精准跨模态匹配与定位。通过协同这些维度,构建视频与文本特征的细粒度特征空间,大幅降低对大规模标注时间标签的需求。
英文摘要:
Video Paragraph Grounding (VPG) aims to precisely locate the most appropriate moments within a video that are relevant to a given textual paragraph query. However, existing methods typically rely on large-scale annotated temporal labels and assume that the correspondence between videos and paragraphs is known. This is impractical in real-world applications, as constructing temporal labels requires significant labor costs, and the correspondence is often unknown. To address this issue, we propose a Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding method (DMR-JRG). In this method, retrieval and grounding tasks are mutually reinforced rather than being treated as separate issues. DMR-JRG mainly consists of two branches: a retrieval branch and a grounding branch. The retrieval branch uses inter-video contrastive learning to roughly align the global features of paragraphs and videos, reducing modality differences and constructing a coarse-grained feature space to break free from the need for correspondence between paragraphs and videos. Additionally, this coarse-grained feature space further facilitates the grounding branch in extracting fine-grained contextual representations. In the grounding branch, we achieve precise cross-modal matching and grounding by exploring the consistency between local, global, and temporal dimensions of video segments and textual paragraphs. By synergizing these dimensions, we construct a fine-grained feature space for video and textual features, greatly reducing the need for large-scale annotated temporal labels.