发表机构
Zhejiang University; Xiaohongshu Inc.(浙江大学; 小红书公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RegRet是一个基于大型多模态模型的区域级检索框架,通过区域感知编码器和多阶段训练增强区域表示,并引入含22.5万对比对的RE GMB基准,在零样本和对比学习后均显著提升检索性能。
AI 中文摘要
区域级检索旨在将用户指定的图像区域与相关区域或文本描述对齐,在电子商务产品搜索和检索增强生成(RAG)等实际应用中发挥着关键作用。尽管近年来大型多模态模型(LMMs)在多模态检索方面取得了显著进展,但它们主要聚焦于全局级任务,难以捕获有效的区域级表示。为弥合这一差距,我们提出了RegRet,一个基于LMM的区域级检索框架,在不损害整体全局检索性能的前提下增强区域表示。其核心在于,RegRet集成了一个区域感知编码器,用于捕获详细的区域特征,同时将其与全局背景上下文进行平衡。为了进一步增强表示的细粒度理解和可区分性,我们设计了一个多阶段训练流程,包括详细的局部描述生成和区域对比学习任务。此外,考虑到当前基准中缺乏区域级对比训练数据以及评估任务多样性有限的问题,我们引入了RE GMB基准。该基准包含22.5万个对比对,覆盖四项多模态检索任务。大量实验验证了我们方法的有效性。在零样本设置下,RegRet优于强基线。进一步结合对比学习训练后,在RE GMB和公开基准上平均提升超过20%,同时在全局级检索任务上取得相当或更好的结果。
英文摘要
Region-level retrieval aims to align user-specified image regions with relevant regions or textual descriptions, playing a crucial role in realworld applications such as e-commerce product search and RAG. Although recent Large Multimodal Models (LMMs) have made significant strides in multimodal retrieval, they primarily focus on global-level tasks and struggle to capture effective region-level representations. To bridge this gap, we present RegRet, an LMM-based Region-level Retrieval framework that enhances the regional representations without compromising overall global retrieval performance. At its core, RegRet integrates a Region-Aware Encoder to capture detailed regional features while balancing them with the global background context. To further enhance the fine-grained understanding and discriminability of representations, we design a multi-stage training pipeline that includes detailed localized captioning and regional contrastive learning tasks. In addition, considering the absence of region-level contrastive training data and the limited diversity of evaluation tasks in current benchmarks, we introduce the REGMB benchmark. It comprises 225k contrastive pairs, covering four multimodal retrieval tasks. Extensive experiments validate the effectiveness of our approach. RegRet outperforms strong baselines in the zero-shot setting. Further training with contrastive learning leads to an average improvement of more than 20\% on both REGMB and public benchmarks, while achieving comparable or better results on global-level retrieval tasks.
CommentsAccepted by ECCV 2026. 22 pages, including references and appendix