arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15320cs.CV

超图正则化的Gramian体积用于多模态检索

Hypergraph-Regularized Gramian Volumes for Multimodal Retrieval

Anindya Nag, Ambuj Mehrish, Sebastiano Vascon

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出超图正则化Gramian体积(HyVol),在训练时利用语义超边增强多模态检索,提升零样本性能,视频到文本检索在MSR-VTT和VATEX上分别提升+8.3和+7.6。

中文摘要 AI 辅助

基于体积的多模态检索联合评分文本查询与候选者的视频、音频和字幕嵌入。虽然这种方法捕捉了候选者内部的高阶对齐,但评分仍局限于候选者本身,语义相关的训练样本主要作为对比负样本。本工作引入了超图正则化Gramian体积(HyVol),一个训练时模块,在评估原始体积损失之前纳入这些语义关系。文档超边连接每个候选者的观测模态,而语义超边连接其分离标题互为top-k邻居的候选者。一个浅层门控超图网络对模态嵌入应用残差校正。存在掩码排除不可用流参与消息传递,身份填充在不进行特征插补的情况下保留观测Gram子矩阵的行列式。由于细化作用于嵌入而非评分,相同的构造适用于Gram和HyperGram。训练后我们移除超图,仅保留骨干架构、原始评分函数和检索成本不变。我们在VAST-27M的150K剪辑子集上训练两个骨干,并在六个基准上评估零样本性能。在配对协议下,HyVol在所有五个检索基准上提升了R@1,视频到文本的增益在MSR-VTT上达到+8.3,在VATEX上达到+7.6。在缺失模态掩码下,V2T边际在所有实验设置中保持正值,尽管T2V边际在四种设置中略微为负。

英文摘要

Volume-based multimodal retrieval jointly scores a text query with a candidate's video, audio, and subtitle embeddings. While this approach captures higher-order within-candidate alignment, the score remains candidate-local, and semantically related training samples primarily serve as contrastive negatives. This work introduces Hypergraph-Regularized Gramian Volumes (HyVol), a training-time module that incorporates these semantic relations prior to evaluating the original volume loss. Document hyperedges connect the observed modalities of each candidate, whereas semantic hyperedges link candidates whose detached captions are mutual top-k neighbors. A shallow gated hyper-graph network applies residual corrections to the modality embeddings. Presence masks exclude unavailable streams from message passing, and identity padding preserves the determinant of the observed Gram submatrix without feature imputation. As refinement operates on embeddings rather than scores, the same construction applies to both Gram and HyperGram. We remove the hypergraph after training, leaving the backbone-only architecture, original scoring function, and retrieval cost unchanged. We train both backbones on a 150K-clip subset of VAST-27M and evaluate zero-shot performance on six benchmarks. Under the paired protocol, HyVol improves R@1 across all five retrieval benchmarks, with video-to-text gains reaching +8.3 on MSR-VTT and +7.6 on VATEX. Under missing-modality masking, the V2T margin remains positive in all experimental settings, although the T2V margin becomes slightly negative in four.

发表机构

  • Ca’ Foscari University of Venice(威尼斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑