arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MulVec:用于无训练零样本组合图像检索的细粒度角色感知匹配

MulVec: Fine-Grained Role-Aware Matching for Training-Free Zero-Shot Composed Image Retrieval

Zihao Zhang, Dayan Wu, Xinze Liu, Hengjie Zhu, Yiliang Zhu, Ding Wang, Peng Fu, Zheng Lin, Weiping Wang

arXiv 2608.25305首次发表:更新:

发表机构

Institute of Information Engineering, CAS; Johns Hopkins University(中国科学院信息工程研究所; 约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MULVEC方法,通过结构化查询的四个检索角色进行细粒度匹配,在三个图像检索数据集及三种骨干规模下提升了零样本组合图像检索性能。

AI 中文摘要

无训练零样本组合图像检索无需从特定任务的图像三元组中学习,即可从图库中的参考图像和文本编辑指令中找到目标图像。现有方法通常将目标作为整体描述,并将该描述与全局图像表示进行匹配,这种全局匹配会混合不同语义线索,丢失细粒度细节。我们提出MULVEC,这是一种角色感知方法,其编译器生成结构化查询记录,映射为四个检索角色:Global(全局)描述完整目标,Desired(期望)说明应出现的内容,Preserve(保留)说明应保留的内容,Forbidden(禁止)说明应消失的内容。冻结的编码器将查询映射为一个目标描述向量和角色特定的探测向量,而每个候选图像由一个全局视觉向量和一组局部视觉向量表示。随后,各检索角色利用此共享证据实现各自目的,通过对其分数进行固定加权求和,在单次检索过程中对整个图库进行排序。在CIRCO、CIRR和FashionIQ三个数据集及三种骨干网络规模下,MULVEC在CIRCO上的mAP@5较最强对比方法提升最高达23.0%,并在对比中取得CIRR和FashionIQ的最佳结果。

英文摘要

Training-free zero-shot composed image retrieval finds a target image in a gallery from a reference image and a text edit without learning from task-specific image triplets. Existing methods typically describe the target as a whole and match this description with a global image representation. This global matching can mix different semantic cues and lose fine- grained details. We propose MULVEC, a role-aware method whose compiler produces a structured query record that is mapped to four retrieval roles: Global describes the full target, Desired states what should appear, Preserve states what should remain, and Forbidden states what should disappear. Frozen encoders map the query to one target description vector and role-specific probe vectors, while each candidate is represented by one global visual vector and a bank of local visual vectors. The retrieval roles then use this shared evidence for their respective purposes, and a fixed weighted sum of their scores ranks the entire gallery in a single retrieval pass. Across CIRCO, CIRR, and FashionIQ and three backbone scales, MULVEC improves CIRCO mAP@5 by up to 23.0% over the strongest compared method and gives the best CIRR and FashionIQ results in our comparison.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑