arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12621cs.CVcs.IR

迈向无视觉的合成图像检索:基于属性增强评分和基于大语言模型的重排的零样本合成图像检索

Towards Vision-Free CIR: Attribute-Augmented Scoring and LLM-Based Reranking for Zero-Shot Composed Image Retrieval

Ryotaro Shimada, Yu-Chieh Lin, Yuji Nozawa, Youyang Ng, Osamu Torii, Yusuke Matsui

首次发表
浏览论文内容

中文总结 AI 辅助

研究如何让无视觉方法有效处理合成图像检索任务,通过属性增强混合评分和基于大语言模型的重排两项关键技术构建框架,在开放域CIRR数据集实验中优于现有零样本方法,在FashionIQ上凸显语义推理与视觉匹配权衡,消融研究验证技术有效性。

中文摘要 AI 辅助

近期工作表明,“无视觉”方法(将图像表示为文本)对标准图像检索任务有效。然而,由于文本描述中存在固有信息损失,该范式能否有效处理更复杂的多模态任务——合成图像检索(CIR)仍不明确。本文介绍了一个无视觉的CIR框架,通过两种关键技术应对这一挑战:(1)属性增强混合评分,通过显式属性匹配补偿丢失的视觉细节;(2)基于大语言模型的重排,验证顶级候选的语义一致性。在开放域CIRR数据集上的实验表明,我们的方法优于现有的零样本CIR方法(R@1提高44.04%,提升8.79%)。在FashionIQ上,结果凸显了语义推理和细粒度视觉匹配之间的权衡。消融研究表明,属性增强评分和基于大语言模型的重排均持续提升性能。

英文摘要

Recent work has shown that "Vision-Free'' approaches (representing images as text) can be effective for standard image retrieval tasks. However, it remains unclear whether this paradigm can effectively handle a more complex, multimodal task, Composed Image Retrieval (CIR), due to the inherent information loss in textual descriptions. In this paper, we introduce a Vision-Free CIR framework that addresses this challenge through two key techniques: (1) Attribute-Augmented Hybrid Scoring, which compensates for lost visual details via explicit attribute matching, and (2) LLM-Based Reranking, which verifies semantic consistency of top candidates. Experiments on the open-domain CIRR dataset show that our approach outperforms existing Zero-shot CIR methods (44.04% R@1, +8.79%). On FashionIQ, our results highlight the trade-off between semantic reasoning and fine-grained visual matching. Ablation studies reveal that both attribute-augmented scoring and LLM-Based Reranking consistently improve performance.

发表机构

  • The University of Tokyo(东京大学)
  • Kioxia Corporation(铠侠株式会社)

机构由 AI 辅助整理,请以论文原文为准。

↑