发表机构
Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出PailitaoGR方法,通过设计聚焦目标感知与选择性辅助证据利用机制,在生成式图像检索任务中较现有基准平均提升13.8%,验证了方法有效性。
AI 中文摘要
生成式检索通过直接生成产品语义标识符(SIDs)已展现出强大性能,但将该范式扩展至图像搜索却并非易事,因为现实查询图像包含多样信息,包括搜索目标、有用辅助证据及无关视觉内容,这要求模型识别并聚焦于搜索目标,同时选择性利用辅助证据。本文提出PailitaoGR,一种用于生成式图像检索的基于图像的潜在思考方法,其将聚焦目标的感知与选择性辅助证据利用内置于生成式检索模型中,实现“缩放不裁剪”与“阅读无需OCR”。具体而言,我们设计了一种聚焦目标的感知机制,由目标增强器和基于在线策略蒸馏与注意力引导损失的学习策略组成,可识别并增强搜索目标的视觉标记,使模型聚焦于搜索目标区域;还设计了一种选择性辅助证据利用机制,由辅助增强器和容量增量对比蒸馏策略组成,可识别并增强辅助证据的视觉标记,使模型能够利用辅助证据。我们从现实在线图像搜索日志中采样构建了训练集与验证集,实验表明,所提方法较现有基准方法平均提升13.8%,验证了其有效性。
英文摘要
Generative retrieval has demonstrated strong performance by directly generating product semantic identifiers (SIDs). Extending this paradigm to image search, however, is nontrivial because real-world query images contain diverse information, including the search target, useful auxiliary evidence, and irrelevant visual content. This requires the model to identify and focus on the search target while selectively utilizing auxiliary evidence. In this paper, we propose \textbf{PailitaoGR}, a \emph{Latent Think-with-Images} method for generative image retrieval, which internalizes target-focused perception and selective auxiliary-evidence utilization into a the generative retrieval model, enabling \textit{Zooming without Cropping} and \textit{Reading without OCR}. Specifically, we design a target-focused perception mechanism that identifies and enhances visual tokens of the search target, consisting of a target Enhancer and a learning strategy based on on-policy distillation and attention guidance loss, enabling the model to focus on search-target regions. We also design a selective auxiliary-evidence utilization mechanism that identifies and enhances visual tokens of auxiliary evidence, including an auxiliary enhancer and an in-capacity incremental contrastive distillation strategy, enabling the model to exploit auxiliary evidence. We construct training and validation sets sampled from real-world online image-search logs. Experiments show that our method outperforms existing baselines by an average of 13.8\%, validating its effectiveness.