发表机构
Shandong University; City University of Hong Kong; Harbin Institute of Technology (Shenzhen); Shandong Jianzhu University(山东大学; 香港城市大学; 哈尔滨工业大学(深圳); 山东建筑大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对MLLMs在复杂图像检索任务中细粒度建模不足的问题,提出自动化细粒度多模态五元组数据集构建流程与两阶段微调策略,在零样本设置下于五个数据集上取得优于现有方法的性能。
AI 中文摘要
凭借强大的通用多模态处理与推理能力,多模态大语言模型(Multimodal Large Language Models, MLLMs)作为通用图像检索器展现出巨大潜力,可有效应对各类真实世界图像检索任务。然而,现有开创性研究虽前景可观,却忽视了细粒度上下文建模与解耦微调目标在提升MLLMs检索性能方面的潜力,尤其在长文本到图像检索、视觉对话检索及组合图像检索(Composed Image Retrieval, CIR)等复杂任务中。因此,本研究提出自动化细粒度多模态五元组数据集构建流程,以及一种新型两阶段细粒度多模态微调策略。该数据集生成流程产出包含细粒度图像描述与修改文本的综合CIR数据集,为细粒度上下文建模提供支持。不同于此前的纠缠式微调范式,本方法将微调过程划分为两个独立阶段:(1)面向细粒度上下文推理的微调;(2)面向细粒度检索的微调。这些阶段旨在依次增强模型的上下文理解能力与查询-目标对齐能力,进而提升检索性能。在涵盖各类复杂图像检索任务的五个数据集上开展的大量实验表明,本方法在零样本检索设置下,相较于现有方法表现出显著优越性,即便采用比这些方法更轻量的MLLM主干模型亦是如此。
英文摘要
Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing diverse real-world image retrieval tasks. Nevertheless, pioneering studies, while promising, overlook the potential of fine-grained context modeling and disentangled fine-tuning objectives in enhancing MLLMs' retrieval performance, particularly for complex tasks such as long-text-to-image retrieval, visual dialog retrieval, and composed image retrieval (CIR). Therefore, in this work, we propose an automated fine-grained multimodal quintuple dataset construction pipeline and a novel two-stage fine-grained multimodal fine-tuning strategy. The dataset generation pipeline produces a comprehensive CIR dataset with fine-grained image captions and modification text, facilitating fine-grained context modeling. Beyond the previously entangled fine-tuning paradigm, our approach separates the fine-tuning process into two distinct stages: (1) fine-grained context reasoning-oriented fine-tuning and (2) fine-grained retrieval-oriented fine-tuning. These stages aim to sequentially enhance the model's context understanding and query-target alignment capabilities, thereby improving retrieval performance. Extensive experiments across five datasets encompassing diverse and complex image retrieval tasks demonstrate the remarkable superiority of our method over existing approaches in zero-shot retrieval settings, even with a more lightweight MLLM backbone compared to those methods.
Journal refProceedings of the ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2025)