arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00551cs.IR

PHA-Net:用于文本-视频检索的基于原型的分层对齐网络

PHA-Net: Prototype-based Hierarchical Alignment Network for Text-Video Retrieval

Xiaolun Jing, Kezhao Yin, Xinxing Yang, Genke Yang, Jian Chu

中文总结 AI 辅助

针对文本-视频检索中跨模态语义不匹配及计算成本高的问题,提出PHA-Net,以模态共享原型为桥梁,结合原型支持的令牌合并模块与原型对比损失,在四个基准数据集上取得显著性能提升。

中文摘要 AI 辅助

随着CLIP等大规模图文预训练模型的出现,文本-视频检索近年来取得了显著进展。现有性能最优的方法需同时在个体、局部和全局层面对齐跨模态语义,引发了对简洁文本与丰富视频之间固有语义不匹配的担忧。一种典型方法是将多个语言-视频注意力模块集成到分层框架中,但该范式仅优化视觉表示,且计算成本高昂。本文提出一种新的基于原型的分层对齐网络(PHA-Net),用于对齐跨模态的个体、局部和全局层面表示。具体而言,我们引入多个模态共享原型作为桥梁,高效优化文本和视频表示以实现跨模态对齐。此外,我们认为聚类令牌中的语义分布不平衡可能会降低检索性能,因为语义薄弱的令牌几乎没有价值。为降低这些令牌的影响,我们提出了原型支持的令牌合并模块,通过原型语义引导增强语义强的令牌并抑制语义弱的令牌。我们还设计了原型对比损失,以鼓励文本和视觉原型聚焦于不同的语义信息,该辅助损失的目标是确保来自同一原型的文本和视觉原型比来自不同原型的原型具有更高的相似度。在四个基准数据集上进行的大量实验证实了PHA-Net的有效性,其在MSR-VTT(所有召回率之和提升8.8%)、ActivityNet(提升19.2%)、VATEX(提升0.7%)和Charades(提升4.9%)上均取得了显著改进。代码可在该https URL获取。

英文摘要

With the emergence of large-scale image-text pre-training models, e.g., CLIP, text-video retrieval has experienced substantial advances in recent years. Existing best-performing methods involve aligning cross-modal semantics at individual, local, and global levels simultaneously, raising concerns about the intrinsic semantic mismatch between concise texts and rich videos. A canonical approach is to integrate multiple language-video attention modules into the hierarchical framework while this paradigm only optimizes visual representations with prohibitive computational costs. In this paper, we propose a new prototype-based hierarchical alignment network (PHA-Net) to align individual/local/global level representations across modalities. Concretely, we introduce multiple modality-shared prototypes as the bridge to efficiently optimize text and video representations for cross-modal alignment. Then, we argue that the imbalanced semantic distribution in clustered tokens may undermine retrieval performance, as tokens with weak semantics are of little interest. To reduce the impact of these tokens, a proposed prototype-supported token merge module is responsible for enhancing tokens with strong semantics and suppressing others with weak semantics via prototype semantics guidance. Moreover, we devise a prototype contrastive loss to encourage textual and visual prototypes to focus on different semantic information. The idea of this auxiliary loss is to ensure higher similarity between textual and visual prototypes from the same prototype than those from different prototypes. Extensive experiments on four benchmarks confirm the effectiveness of our PHA-Net, which achieves significant improvements in the sum of all recalls on MSR-VTT (8.8%), ActivityNet (19.2%), VATEX (0.7%), and Charades (4.9%). Code is available at https://github.com/JingXiaolun/PHA-Net.

↑