arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25412cs.CV

AdaptiveEmbed:面向多模态检索的样本自适应多向量表示

AdaptiveEmbed: Sample-Adaptive Multi-Vector Representation for Multimodal Retrieval

Xinze Liu, Lei Yang, Dayan Wu, Hengjie Zhu, Zihao Zhang, Hanqi Wu, Tianzhu Hu, Peng Fu, Zheng Lin, Weiping Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出样本自适应多向量表示(SAMVR)新设定,构建AdaptiveEmbed框架,通过多组对比学习等技术实现样本级自适应容量分配,在多模态检索基准上验证其性能优于固定容量方案。

中文摘要 AI 辅助

多向量表示已成为多模态检索领域的有效范式,它用多个互补嵌入来表示每个样本,以捕捉细粒度的跨模态信息。然而,现有方法通常采用固定的表示容量,给所有样本分配相同数量的向量,而不考虑它们各自的检索需求。这种固定容量的设定忽略了一个事实:不同样本可能需要不同的表示容量才能实现有效检索。在本研究中,我们提出了样本自适应多向量表示(Sample-Adaptive Multi-Vector Representation, SAMVR),这是一种用于多模态检索的新问题设定,旨在研究如何在样本层面分配多向量表示容量。在SAMVR框架下,每个样本由内容自适应嵌入集(Content-Adaptive Embedding Set, CAES)表示,其容量根据额外表示向量的样本特定检索效用确定。为了实例化SAMVR,我们提出了AdaptiveEmbed,这是一个用于学习样本自适应多向量表示的统一框架。AdaptiveEmbed通过多组对比学习(Multi-Group Contrastive Learning, MGCL)结合对称的集对集相似度(set-to-set similarity, SetSim)来学习结构化多向量表示,还采用效用策略优化(Utility Policy Optimization, UPO),通过边际效用分配(Marginal Utility Allocation, MUA)来确定样本特定的表示容量。在涉及图像、文本、视频和音频的多模态检索基准上开展的实验表明,样本自适应容量分配相比固定容量的多向量表示实现了更优的整体检索性能,验证了SAMVR在多模态检索中的有效性。这些结果确立了SAMVR作为多向量多模态检索中自适应容量分配的可行设定。

英文摘要

Multi-vector representations have emerged as an effective paradigm for multimodal retrieval, representing each sample with multiple complementary embeddings to capture fine-grained cross-modal information. However, existing approaches typically employ a fixed representation capacity, assigning the same number of vectors to all samples regardless of their individual retrieval demands. Such a fixed-capacity formulation overlooks the fact that different samples may require different amounts of representation capacity for effective retrieval. In this work, we introduce \emph{Sample-Adaptive Multi-Vector Representation} (SAMVR), a new problem setting for multimodal retrieval that studies how multi-vector representation capacity can be allocated at the sample level. Under SAMVR, each sample is represented by a \emph{content-adaptive embedding set} (CAES), whose capacity is determined according to the sample-specific retrieval utility of additional representation vectors. To instantiate SAMVR, we propose \emph{AdaptiveEmbed}, a unified framework for learning sample-adaptive multi-vector representations. AdaptiveEmbed learns structured multi-vector representations through \emph{Multi-Group Contrastive Learning} (MGCL) with the symmetric \emph{set-to-set similarity} (SetSim), and further employs \emph{Utility Policy Optimization} (UPO) to determine sample-specific representation capacity via \emph{Marginal Utility Allocation} (MUA). Experiments across multimodal retrieval benchmarks involving image, text, video, and audio show that sample-adaptive capacity allocation achieves overall better retrieval performance than fixed-capacity multi-vector representations, validating the effectiveness of SAMVR for multimodal retrieval. These results establish SAMVR as a viable formulation for adaptive capacity allocation in multi-vector multimodal retrieval.

发表机构

  • Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
  • School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)

机构由 AI 辅助整理,请以论文原文为准。

↑