arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

弥合目录与现实的差距:通过多阶段对比学习实现可扩展的产品识别

Bridging the Catalog-to-Real Gap: Scalable Product Recognition via Multi-Stage Contrastive Learning

Anyi Zhang, Joy Mazumder, Kiril Lomakin

arXiv 2607.09888首次发表:更新:

发表机构

Walmart Global Tech(沃尔玛全球技术公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自动产品识别中现实与目录匹配的可扩展性瓶颈,提出将任务转为跨域检索问题,引入目录到现实多阶段对比学习范式,利用项目和图像级相似性微调视觉主干,能无缝扩展到未见产品和类别,有出色零样本泛化性能。

AI 中文摘要

自动产品识别是现代零售智能的基石,但将现实世界中的店内图像与大量公司目录准确匹配,仍是大规模应用的主要可扩展性瓶颈。本文将该任务重新表述为基于嵌入的跨域检索问题,而非标准的封闭集分类任务。具体而言,目标是从大量库存中为给定的现实世界产品查询裁剪检索最匹配的目录参考图像。为弥合原始工作室包装照片与嘈杂店内查询之间的严重域差距,引入了一种新颖的目录到现实多阶段对比学习范式(Cat2Real)。该框架通过系统利用项目级和图像级相似性来微调视觉主干,以驱动有针对性的硬负样本挖掘。大量实证评估表明,该范式可无缝扩展到未见产品和类别,即使在完全没有针对新库存的现实世界训练图像的情况下,也能产生出色的零样本泛化性能。

英文摘要

Automated product recognition is a cornerstone of modern retail intelligence; however, accurately matching real-world, in-store images against extensive corporate catalogs remains a major scalability bottleneck for large-scale applications. In this work, we address this challenge by reformulating the task as an embedding-based cross-domain retrieval problem rather than a standard closed-set classification task. Specifically, we define the objective as retrieving the most corresponding catalog reference image for a given real-world product query crop from an expansive inventory. To bridge the severe domain gap between pristine studio packshots and noisy in-store queries, we introduce a novel catalog-to-real multi-stage contrastive learning paradigm (Cat2Real). This framework fine-tunes a vision backbone by systematically exploiting both item-level and image-level similarities to drive targeted hard negative mining. Extensive empirical evaluations demonstrate that our paradigm scales seamlessly to unseen products and categories, yielding outstanding zero-shot generalization performance even in the complete absence of real-world training images for novel inventory.

Comments12 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑