arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15335cs.CV

查询条件下的球面质心聚合用于多模态检索

Query-Conditioned Spherical Centroid Aggregation for Multimodal Retrieval

Ambuj Mehrish, Anindya Nag, Sebastiano Vascon

AI总结:

提出查询条件下的SCALAR聚合器,通过自适应权重融合多模态,在多个基准上超越对称聚合方法,提升检索性能。

AI中文摘要:

多模态检索整合了视频、音频、字幕和文本;然而,最近的几何聚合器,如Gramian体积、双曲体积和谱目标,将所有模态对称对待。在统一评估协议下,它们的联合分数经常落后于最强的单模态路径1.9至27.6 R@1。受控分析将此结果归因于均匀的模态影响。本工作引入了具有学习自适应相关性的球面质心聚合(SCALAR),这是一种查询条件下的聚合器,在计算球面质心之前为每个可用模态分配基于相关性的权重。SCALAR适应任意模态子集,并使用秩为8的LoRA适配器在掩蔽的、减少元数的视图上进行训练。在五个基准上,SCALAR在四个基准上实现了正的聚合增益,达到+4.0 R@1,而所评估的先前的聚合器没有一个在超过一个基准上为正。均匀权重的消融重现了对称聚合中观察到的退化。仅用480万个可训练参数,SCALAR在三个基准上取得了最高的文本到视频R@1,并在第四个基准上在种子变异范围内达到最佳结果。在测试时模态丢弃下,SCALAR在表示阶段的分数在每个评估的掩蔽率和基准上都超过了发布的GRAM检查点,高出3.2至10.9 R@1。最后,随着模态被移除,仅在完整模态集上训练的重新排序器越来越趋向于它们的仅视频路径,减少了这些表示级别的增益,并强调了标准两阶段检索流程的局限性。

英文摘要:

Multimodal retrieval integrates video, audio, subtitles, and text; however, recent geometric aggregators, such as Gramian volumes, hyperbolic volumes, and spectral objectives, treat all modalities symmetrically. Under a unified evaluation protocol, their joint scores frequently lag behind the strongest single-modality pathway by 1.9 to 27.6 R@1. Controlled analyses attribute this outcome to uniform modality influence. This work introduces Spherical Centroid Aggregation with Learned Adaptive Relevance (SCALAR), a query-conditioned aggregator that assigns relevance-based weights to each available modality before computing a spherical centroid. SCALAR accommodates arbitrary modality subsets and is trained on masked, reduced-arity views using rank-8 LoRA adapters. Across five benchmarks, SCALAR achieves positive aggregation gain on four, reaching +4.0 R@1, while none of the evaluated prior aggregators is positive on more than one. A uniform-weight ablation reproduces the degradation observed with symmetric aggregation. With only 4.8 million trainable parameters, SCALAR attains the highest text-to-video R@1 on three and performs within seed variation of the best result on a fourth. Under test-time modality dropout, SCALAR's representation-stage score surpasses the released GRAM checkpoint at every evaluated masking rate and benchmark by 3.2 to 10.9 R@1. Finally, as modalities are removed, rerankers trained exclusively on complete modality sets increasingly converge toward their video-only pathways, diminishing these representation-level gains and underscoring a limitation of standard two-stage retrieval pipelines.

↑