发表机构
Walmart Global Tech(沃尔玛全球技术公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自动产品识别中现实与目录匹配的可扩展性瓶颈,提出将任务转为跨域检索问题,引入目录到现实多阶段对比学习范式,利用项目和图像级相似性微调视觉主干,能无缝扩展到未见产品和类别,有出色零样本泛化性能。
AI 中文摘要
自动产品识别是现代零售智能的基石,但将现实世界中的店内图像与大量公司目录准确匹配,仍是大规模应用的主要可扩展性瓶颈。本文将该任务重新表述为基于嵌入的跨域检索问题,而非标准的封闭集分类任务。具体而言,目标是从大量库存中为给定的现实世界产品查询裁剪检索最匹配的目录参考图像。为弥合原始工作室包装照片与嘈杂店内查询之间的严重域差距,引入了一种新颖的目录到现实多阶段对比学习范式(Cat2Real)。该框架通过系统利用项目级和图像级相似性来微调视觉主干,以驱动有针对性的硬负样本挖掘。大量实证评估表明,该范式可无缝扩展到未见产品和类别,即使在完全没有针对新库存的现实世界训练图像的情况下,也能产生出色的零样本泛化性能。
英文摘要
Automated product recognition is a cornerstone of modern retail intelligence; however, accurately matching real-world, in-store images against extensive corporate catalogs remains a major scalability bottleneck for large-scale applications. In this work, we address this challenge by reformulating the task as an embedding-based cross-domain retrieval problem rather than a standard closed-set classification task. Specifically, we define the objective as retrieving the most corresponding catalog reference image for a given real-world product query crop from an expansive inventory. To bridge the severe domain gap between pristine studio packshots and noisy in-store queries, we introduce a novel catalog-to-real multi-stage contrastive learning paradigm (Cat2Real). This framework fine-tunes a vision backbone by systematically exploiting both item-level and image-level similarities to drive targeted hard negative mining. Extensive empirical evaluations demonstrate that our paradigm scales seamlessly to unseen products and categories, yielding outstanding zero-shot generalization performance even in the complete absence of real-world training images for novel inventory.
Comments12 pages, 4 figures