arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25037cs.AIcs.CLcs.DBcs.IR

检索、匹配、升级:基于VLM蒸馏的交叉编码器与智能体VLMs的准确且可扩展的商品链接

Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs

Jian Wang, Steven Xu, Sanjyot Thete, Maryam Barouti, Tom Tang, Elaine Wu, Charu Sareen, Kyle MacDonald

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对大规模商品链接的成本与精度矛盾,提出“检索-匹配-升级”级联框架,用VLM蒸馏的交叉编码器处理多数简单案例,智能体VLM解决模糊案例,提升了链接覆盖率且降低成本。

中文摘要 AI 辅助

商品链接是将商家商品记录映射到规范目录商品的实体对齐任务,用于整合碎片化列表,使下游搜索、推荐和广告能为每个商品提供一个清晰条目。在市场规模下,需将数十亿条嘈杂的多类别记录与数千万条规范商品进行对齐,此时用单个模型对每个候选打分要么对困难案例太弱,要么对简单案例成本过高。本文提出一种生产级的“先检索后匹配”级联框架,按难度分配计算资源:检索环节生成合理匹配项,轻量文本交叉编码器自动解决高置信度的多数情况,智能体多模态视觉语言模型(agentic VLM)通过检查商品图像并对两个记录都未包含的证据进行网页搜索,解决模糊的剩余案例。该交叉编码器由数百万个双VLM共识标签蒸馏而来,训练集摒弃了人工标注,并经校准以在操作员认证的小型审计验证的98%精度阈值下自动接受链接。该智能体是自托管的开放权重模型,在仅损失4个点召回率(88% vs 92%)的情况下达到了封闭前沿VLM的精度,且每对成本仅为其约1/7,无需微调。每对成本从廉价交叉编码器到前沿VLM跨度近5个数量级,因此仅将困难尾部升级到智能体,可将端到端链接覆盖率从廉价阶段的68%提升至77%。

英文摘要

Product linking, the entity-resolution task of mapping merchant product records to canonical catalog products, consolidates fragmented listings so downstream search, recommendation, and advertising see one clean entry per product. At marketplace scale, billions of noisy, multi-category records must be resolved against tens of millions of canonical products, where scoring every candidate with a single model is either too weak for the hard cases or too costly for the easy ones. We present a production retrieve-then-match cascade that spends computation in proportion to difficulty: retrieval surfaces plausible matches, a lightweight text cross-encoder auto-resolves the high-confidence majority, and an agentic multimodal vision-language model settles the ambiguous remainder by inspecting product images and issuing web searches for evidence that is in neither record. The cross-encoder is distilled from millions of dual-VLM-consensus labels, retiring human annotation from the training set, and is calibrated to auto-accept links at a 98% precision bar validated against a smaller operator-certified audit. The agent is a self-hosted open-weight model that reaches a closed frontier VLM's precision at a four-point recall cost (88% versus 92%) for roughly one-seventh the per-pair cost, with no fine-tuning. Per-pair cost spans nearly five orders of magnitude from the cheap cross-encoder to the frontier VLM, so escalating only the hard tail to the agent raises end-to-end link coverage from the cheap stage's 68% to 77%.

发表机构

  • DoorDash Inc.(DoorDash公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑