保留与组合训练用于组合图像检索
Preserve-and-Compose Training for Composed Image Retrieval
浏览论文内容
中文总结 AI 辅助
针对组合图像检索中目标标题遗漏源细节的问题,提出PACT训练方法,利用源图像视觉证据补充监督,并引入Chord评分,在四个基准上显著提升零样本检索性能。
中文摘要 AI 辅助
组合图像检索(CIR)旨在检索满足用户指定修改的图像,同时保留参考图像中相关的视觉内容。为此收集目标图像成本高昂,促使了零样本CIR方法使用目标标题作为监督信号。然而,目标标题可能省略应保留的源细节。因此,我们提出保留与组合训练,该方法通过源图像的视觉证据补充目标标题监督。PACT从图像-文本-文本(ITT)三元组中学习,无需目标图像或图库更新,将组合查询与目标标题对齐,同时通过视觉监督保留源证据。我们进一步引入Chord评分,在冻结图像空间中结合目标相似性与源相对方向一致性。在四个ZS-CIR基准上的结果表明,将目标标题监督与源图像证据相结合,能在不同数据集、骨干网络规模和外部图库上实现强检索性能。代码可在该https URL获取。
英文摘要
Composed image retrieval (CIR) aims to retrieve images that satisfy a user-specified modification while preserving relevant visual content from a reference image. Collecting target images for this purpose is costly, motivating zero-shot CIR methods that instead use target captions as supervision. However, target captions may omit source details that should be preserved. We therefore propose, Preserve-and-Compose Training, which complements target-caption supervision with visual evidence from the source image. PACT learns from image--text--text (ITT) triplets without target images or gallery updates, aligning composed queries with target captions while preserving source evidence through visual supervision. We further introduce Chord scoring, which combines target similarity with source-relative directional agreement in the frozen image space. Results across four ZS-CIR benchmarks show that combining target-caption supervision with source-image evidence leads to strong retrieval performance across datasets, backbone scales, and external galleries. The code is available on https://github.com/sehyunkwon/PACT.
发表机构
- Hanyang University ERICA(汉阳大学ERICA校区)
机构由 AI 辅助整理,请以论文原文为准。