arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

优化VLP对齐的多模态意图表示与正确视觉实例化用于零样本组合图像检索

Optimizing VLP-aligned Multimodal Intent Representation with Correct Visual Instantiation for Zero-Shot Composed Image Retrieval

Xuri Ge, Chunhao Wang, Junchen Fu, Haokun Wen, Zhiwei Xu, Ying Zhou, Zhumin Chen, Pengjie Ren, Zhaochun Ren, Xin Xin

arXiv 2609.36946首次发表:更新:

发表机构

School of Artificial Intelligence, Shandong University; School of Computer Science and Technology, Shandong University; School of Computing Science, University of Glasgow; School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen); Leiden University(山东大学人工智能学院; 山东大学计算机科学与技术学院; 格拉斯哥大学计算科学学院; 哈尔滨工业大学(深圳)计算机科学与技术学院; 莱顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出VMIR-CVI框架,通过VLP对齐的文本意图推理与无需训练的视觉实例解耦,优化多模态意图表示,在三个基准上实现零样本组合图像检索的最优性能。

AI 中文摘要

零样本组合图像检索(ZS-CIR)旨在无需配对监督的情况下,从参考图像和修改文本中检索目标图像,通常通过将组合查询编码为视觉-语言预训练模型(VLP)图像-文本匹配空间中的文本主导表示来实现。然而,通过视觉伪词学习或基于多模态大语言模型(MLLM)的目标推理重建的查询,常常偏离VLP原生表示空间,前者因参考噪声和粗糙的文本融合,后者因冗长且视觉基础薄弱的描述。本文提出一个统一的ZS-CIR框架(名为VMIR-CVI),从两个互补视角重建多模态组合查询,以优化VLP兼容的多模态意图表示。首先,它将多模态意图推理并转换为统一的文本描述,与VLP骨干网络的原生文本空间对齐,生成更具检索兼容性的文本查询。其次,它利用正确解耦的视觉实例线索重建查询表示,减少参考噪声同时保留目标相关内容。具体而言,一个VLP对齐的多模态意图推理(VMIR)模块将少样本VLP风格示例注入思维链提示中,引导MLLM生成与目标一致的意图查询。一个无需训练的视觉实例解耦(TVID)模块在不进行额外优化的情况下,从全局参考特征中解耦细粒度视觉实例。最后,一个轻量级的混合模态意图对齐与融合(HIAF)模块将推理的文本意图和解耦的视觉线索整合为统一的混合模态表示,以实现稳健的ZS-CIR。在三个CIR基准(即CIRR、CIRCO和FashionIQ)上的大量实验表明,VMIR-CVI显著优于现有基线,并取得了新的最先进性能。代码和训练模型将公开发布。

英文摘要

ZS-CIR aims to retrieve a target image from a reference image and a modification text without paired supervision, typically by encoding composed queries as text-dominant representations within the image-text matching space of VLPs. However, queries reconstructed by visual pseudo-word learning or MLLM-based target reasoning often deviate from the native VLP representation space due to reference noise and coarse text fusion in the former, and verbose, weakly visually grounded descriptions in the latter. In this paper, we propose a unified ZS-CIR framework (named VMIR-CVI) to reconstruct multimodal composite queries from two complementary perspectives for optimizing VLP-compatible multimodal intent representation. First, it reasons and converts the multimodal intent into a unified textual description, aligning with the native text space of the VLP backbones to produce more retrieval-compatible textual queries. Second, it reconstructs the query representation with correctly decoupled visual instance cues, reducing reference noise while preserving target-relevant content. Specifically, a VLP-aligned Multimodal Intent Reasoning (VMIR) module injects few-shot VLP-style exemplars into chain-of-thought prompts, guiding the MLLM to generate target-consistent intent queries. A Training-free Visual Instance Disentanglement (TVID) module decouples fine-grained visual instances from global reference features without additional optimization. Finally, a lightweight Hybrid-modal Intent Alignment and Fusion (HIAF) module integrates the reasoned textual intent and disentangled visual cues into a unified hybrid-modal representation for robust ZS-CIR. Extensive experiments on three CIR benchmarks, namely CIRR, CIRCO and FashionIQ, show that VMIR-CVI significantly outperforms existing baselines and achieves new state-of-the-art performance. Code and trained models will be publicly released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑