arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37225cs.CVcs.AI

ResComEmb:通过残差同质性压缩实现高效的多模态嵌入

ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression

Zijing Cai, Yuzhe Wang, Jingxian Zhu, Fengbin Zhu, Richang Hong

首次发表
浏览论文内容

中文总结 AI 辅助

ResComEmb提出一种可训练的多向量多模态嵌入框架,通过残差同质性压缩减少冗余,并采用长度自适应双向后期交互匹配,在低视觉令牌预算下实现优于现有方法的检索性能。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)在通用多模态表示学习方面展现出强大潜力。然而,现有方法要么将每个输入压缩为单个向量,限制了细粒度的表达能力;要么保留长序列的视觉令牌向量,导致大量的存储和交互成本。为解决这一权衡问题,我们提出了ResComEmb,一个可训练的框架,用于实现有效且高效的通用多向量多模态嵌入。ResComEmb首先将每个输入以原生动态分辨率编码为有序的全局、中间和细粒度视图。经过MLLM上下文化和嵌入投影后,一个可训练的残差同质性压缩(RHC)模块在显式视觉令牌预算下减少粒度内冗余和跨粒度重复。然后,ResComEmb引入了一种长度自适应的双向后期交互匹配机制,用于稳健的查询-文档评分,该机制在每一方向上平均最强的令牌级匹配,并根据两侧各自拥有的有效令牌数量计算权重来结合两个分数。在MMEB、ViDoRe V1和ViDoRe V2上的大量实验表明,ResComEmb生成的高质量通用多模态嵌入优于VLM2Vec-V2,并且在视觉文档检索中仅使用ColQwen2.5完整视觉令牌预算的37.5%即超越后者,展示了良好的效果-效率权衡。

英文摘要

Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Hefei University of Technology(合肥工业大学)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑