ResComEmb:通过残差同质性压缩实现高效的多模态嵌入
ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression
浏览论文内容
中文总结 AI 辅助
ResComEmb提出一种可训练的多向量多模态嵌入框架,通过残差同质性压缩减少冗余,并采用长度自适应双向后期交互匹配,在低视觉令牌预算下实现优于现有方法的检索性能。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)在通用多模态表示学习方面展现出强大潜力。然而,现有方法要么将每个输入压缩为单个向量,限制了细粒度的表达能力;要么保留长序列的视觉令牌向量,导致大量的存储和交互成本。为解决这一权衡问题,我们提出了ResComEmb,一个可训练的框架,用于实现有效且高效的通用多向量多模态嵌入。ResComEmb首先将每个输入以原生动态分辨率编码为有序的全局、中间和细粒度视图。经过MLLM上下文化和嵌入投影后,一个可训练的残差同质性压缩(RHC)模块在显式视觉令牌预算下减少粒度内冗余和跨粒度重复。然后,ResComEmb引入了一种长度自适应的双向后期交互匹配机制,用于稳健的查询-文档评分,该机制在每一方向上平均最强的令牌级匹配,并根据两侧各自拥有的有效令牌数量计算权重来结合两个分数。在MMEB、ViDoRe V1和ViDoRe V2上的大量实验表明,ResComEmb生成的高质量通用多模态嵌入优于VLM2Vec-V2,并且在视觉文档检索中仅使用ColQwen2.5完整视觉令牌预算的37.5%即超越后者,展示了良好的效果-效率权衡。
英文摘要
Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- Hefei University of Technology(合肥工业大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。