文本-视频检索:基于多维显著性评估与粒度感知查询分解
Text-Video Retrieval via Multi-Dimensional Saliency Assessment and Granularity-Aware Query Decomposition
浏览论文内容
中文总结 AI 辅助
提出MMTI方法,通过多维显著性评估与关键特征选择缓解视觉冗余,并利用动态门控将文本查询分解为句子、帧和补丁查询,实现多粒度文本-视频对齐,在四个基准上超越现有方法。
中文摘要 AI 辅助
文本-视频检索旨在通过学习联合嵌入空间来桥接视觉和文本模态,已成为多模态智能中的关键任务。尽管已有大量工作致力于缓解视觉冗余,以往方法通常依赖单一方面的标准来评估视觉重要性,忽视了视频多方面的时空特性。此外,将文本编码为单一全局嵌入以与视频对齐,会将时间事件和空间实体压缩到统一表示空间中,进一步加剧跨模态错配。为解决这些问题,我们提出MMTI方法,该方法联合缓解视觉冗余并实现多粒度文本-视频交互,以达到准确的多粒度语义对齐。具体而言,关键特征选择(KFS)机制通过联合评估多维显著性和可学习的重要性分数,自适应地识别并聚合信息丰富的帧和补丁,有效压缩密集视觉特征并缓解视觉冗余。此外,我们提出的多粒度文本-视频交互模块(TVIM)采用动态门控机制,将文本查询分解为句子、帧和补丁查询(SFP),实现多粒度文本-视频对齐。由此实现了不同粒度上的互补对齐。在四个标准基准上的大量实验表明,我们的方法优于现有最先进方法。
英文摘要
Text-video retrieval, which aims to bridge visual and textual modalities by learning a joint embedding space, has become a crucial task in multimodal intelligence. Despite extensive efforts to mitigate visual redundancy, previous methods typically rely on a single-aspect criterion to assess visual importance, overlooking the multifaceted spatiotemporal nature of video. In addition, encoding text into a single global embedding to align with videos compresses temporal events and spatial entities into a unified representation space, further aggravating cross-modal misalignment. To address these issues, we propose MMTI, a method that jointly mitigates visual redundancy and enables multi-grained text-video interaction to achieve accurate multi-grained semantic alignment. Specifically, a key feature selection (KFS) mechanism adaptively identifies and aggregates informative frames and patches by jointly evaluating multi-dimensional saliency and learnable importance scores, effectively compacting dense visual features and mitigating visual redundancy. Furthermore, our proposed multi-grained text-video interaction module (TVIM) employs a dynamic gating mechanism to decompose the text query into sentence, frame, and patch queries (SFP), enabling multi-grained text-video alignment. Complementary alignment at different granularities is thereby achieved. Extensive experiments on four standard benchmarks demonstrate that our method outperforms state-of-the-art methods.