发表机构
University of Luxembourg(卢森堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出无需训练的PeFuse伪融合框架,结合扩散模型与多模态大语言模型,将组合图像检索转化为单模态检索任务,在零样本场景下实现了媲美当前最优方法的性能。
AI 中文摘要
组合图像检索(Composed Image Retrieval,CIR)是基于内容的图像检索领域的新兴范式,支持用户通过参考图像与辅助模态(通常为文本)结合的方式构建组合查询,可实现细粒度搜索——目标图像与用户提供的参考图像共享结构元素,同时融入辅助文本指定的修改内容。传统CIR方法依赖多模态融合将视觉与文本特征合并为联合查询嵌入,需训练模块对齐组合查询与目标图像。本文提出PeFuse(伪融合,pseudo-fusion),一种无需训练的框架,利用预训练的Diffusion Models(扩散模型)和Multimodal Large Language Models(多模态大语言模型)通过生成式转换桥接模态,引入单向与双向转换两种新策略,将CIR转化为四类单模态检索问题,重定义为模态内或跨模态单查询检索任务,规避专用任务特定训练需求。在标准基准上的大量实验表明,将CIR转换为文本到图像检索任务比其他转换策略更有效,性能优于或媲美当前最优方法,且因转换流程组件可替换保持高灵活性,验证了伪融合范式在零样本CIR中的有效性,代码公开于:this https URL。
英文摘要
Composed Image Retrieval (CIR) is an emerging paradigm in content-based image retrieval that enables users to formulate compositional queries by combining a reference image with an auxiliary modality, usually text-based. This approach supports fine-grained search where the target image shares structural elements with the user-provided image while incorporating the modifications specified by the auxiliary text. Conventional CIR methods rely on multimodal fusion to combine visual and textual features into a joint query embedding, which requires training modules that align composed queries with the targets. In this work, we propose PeFuse (for pseudo-fusion), a training-free framework that leverages pretrained Diffusion Models and Multimodal Large Language Models to bridge modalities via generative conversion. We introduce two novel strategies: uni-directional and bi-directional conversion, which convert CIR into four single-modality retrieval problems. These methods reformulate CIR as either intra-modal or cross-modal single-query retrieval tasks, bypassing the need for dedicated task-specific training. Extensive experiments on standard benchmarks demonstrate that converting CIR into text-to-image retrieval tasks is more effective than alternative conversion strategies, achieving competitive or superior performance compared with state-of-the-art methods, while maintaining high flexibility thanks to replaceable components of the conversion pipeline. These results highlight the effectiveness of the pseudo-fusion paradigm for zero-shot CIR. Our code is publicly available at: https://github.com/StevenXuf/PeFuse4CIR.
Journal refTransactions on Machine Learning Research, 2026