arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoCo-IR:上下文组合图像检索

CoCo-IR: Contextual Composed Image Retrieval

Shengcao Cao, Tanmaya Shekhar Dabral, Zhongli Ding, Madhuri Shanbhogue, Kaifeng Chen, Zhe Li, Mojtaba Seyedhosseini, Yu-Xiong Wang, Liang-Yan Gui

arXiv 2608.05149首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Google DeepMind; OpenAI(伊利诺伊大学厄巴纳-香槟分校; 谷歌DeepMind; OpenAI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有图像检索系统无法处理多轮交互的局限,提出CoCo-IR任务,构建基于LMM的模型并开发自主数据引擎,在CIRCO和自建CoCo-IR基准上取得最优性能。

AI 中文摘要

当前基于指令的图像检索系统功能强大,但仅限于单轮交互,无法捕捉复杂现实视觉搜索的迭代特性。为克服这一局限,我们提出上下文组合图像检索(CoCo-IR)这一新任务,使用户可通过交互逐步优化搜索结果。我们基于多模态大模型(LMM)构建新模型,作为CoCo-IR的上下文感知推理器,该模型会解析全部交互历史,生成随轮次演化的可变换图像嵌入(TIE)。为在无需昂贵人工标注的情况下支撑模型训练,我们开发了一套完全自主、可扩展的数据引擎,利用LMM生成高质量上下文检索数据,并通过模型引导的验证挖掘具有挑战性的难负样本。大量实验表明,我们的方法达到了新的最优性能:在极具挑战性的单轮基准CIRCO上,我们取得了39.4的mAP@5;此外,在我们的新CoCo-IR基准上,模型在4轮对话中保持了44.1的R@1,显著优于现有方法(4轮R@1为28.2),现有方法无法处理多轮上下文。项目页面:this https URL。

英文摘要

Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of complex, real-world visual searches. To overcome this limitation, we introduce Contextual Composed Image Retrieval (CoCo-IR), a novel task that enables users to progressively refine search results through interactions. We address this new task by proposing a new model based on a Large Multimodal Model (LMM) that functions as a context-aware reasoner for CoCo-IR. Our model interprets the entire interaction history to generate Transformable Image Embeddings (TIE) that evolve across turns. To fuel the model training without expensive human annotations, we develop a fully autonomous, scalable data engine that leverages LMMs to generate high-quality contextual retrieval data, and uses model-guided verification to mine challenging hard negatives. Extensive experiments demonstrate that our approach establishes new state-of-the-art performance: We achieve 39.4 mAP@5 on the challenging single-turn benchmark CIRCO; furthermore, on our new CoCo-IR benchmark, our model maintains robust performance with 44.1 R@1 on 4-turn dialogues, dramatically outperforming existing methods (28.2 4-turn R@1) that fail to handle multi-turn context. Project page: https://CoCo-IR.github.io.

CommentsECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑