发表机构
Walmart Global Tech(沃尔玛全球技术)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对电商视觉搜索中组合图像检索的相关性二元假设问题,提出GradCIR方法,利用VLM生成分级标签、难负样本挖掘和层次感知目标训练检索器,在NDCG@10上提升达5.9%,并已部署于沃尔玛生产系统。
AI 中文摘要
大规模电商目录上的视觉搜索必须同时服务于两类查询:“相似性”查询,即寻找与上传图像相似的物品;以及“修改器”查询,即由图像和描述所需修改(例如颜色变化或风格更换)的文本组成。后者即组合图像检索(CIR)的设置。然而,现有的CIR方法将相关性视为二元的,并在具有单一正样本的三元组上训练——这并不适合真实目录,因为真实目录中许多候选物品部分满足用户查询,而跨该部分匹配谱系的排序驱动着客户体验。我们提出了一种在分级相关性上训练CIR检索器的方法,包括:(i)使用视觉语言模型(VLM)来整理训练数据,生成查询(目标检测+修改器合成)和4级相关性标签,无需人工标注;(ii)一个迭代的相关性反馈循环,通过从训练中的检索器中挖掘难负样本来扩展训练集;(iii)一个层次感知的角度目标函数,直接基于分级标签训练检索器,而不是将其折叠为二元分割。我们将此方法命名为GradCIR,并在基于从原始沃尔玛目录数据整理的350万分级对训练的PaliGemma2双编码器上实例化。一项受控的分级与二元消融实验隔离了监督粒度,并显示NDCG@10提升了4.9%-5.9%。相同的方案应用于其他多模态编码器,将早期融合骨干网络的NDCG@10提升高达8.5%。在公开的FashionIQ基准上,GradCIR(应用于PaliGemma2)在微调后达到0.6703的平均召回率,略高于我们比较的最强同行评审监督基线,并匹配或超过所有已发表的CLIP-L类零样本CIR方法。该系统已在沃尔玛生产环境中部署,正在服务实时视觉搜索用户流量。
英文摘要
Visual search on large e-commerce catalogs must serve both "similarity" queries that ask for items resembling an uploaded image and "modifier" queries that comprise an image and text describing a desired modification (e.g. a color change or style swap). The latter is the setting known as composed image retrieval (CIR). Existing CIR methods, however, treat relevance as binary and train on triplets with a single positive target - a poor fit for real catalogs where many candidates partially satisfy a user query and ranking across that partial-match spectrum drives the customer experience. We propose a methodology for training CIR retrievers on graded relevance, consisting of: (i) a VLM to curate training data, generating both queries (object detection + modifier synthesis) and 4-level relevance labels without manual annotation, (ii) an iterative relevance-feedback loop that expands the training set by mining hard negatives from the in-training retriever, and (iii) a hierarchy-aware angular objective to train the retriever directly on the graded labels rather than collapsing them to a binary split. We call this methodology GradCIR and instantiate it on a PaliGemma2 bi-encoder trained on 3.5M graded pairs curated from raw Walmart catalog data. A controlled graded-vs-binary ablation isolates the supervision granularity and shows lift of 4.9%-5.9% in NDCG@10. The same recipe applied to other multimodal encoders lifts early-fusion backbones by up to 8.5% NDCG@10. On the public FashionIQ benchmark, GradCIR (applied to PaliGemma2) reaches 0.6703 average recall when fine-tuned, slightly ahead of the strongest peer-reviewed supervised baseline we compare against, and matching or exceeding all published CLIP-L-class zero-shot CIR methods. The system is deployed in production at Walmart, where it's serving live visual-search user traffic.