AI 中文总结
研究基于嵌入的检索在大规模电子商务搜索中的问题,提出统一管道,采用混合硬负样本挖掘和遗留感知蒸馏方法,经实验验证投入生产,提升了NDCG@5和总收入。
AI 中文摘要
基于嵌入的检索(EBR)是大规模电子商务搜索的基础,但其有效性常受训练信号质量和编码器表示能力的限制。标准双编码器存在训练-推理差距,向更高容量主干过渡会导致检索行为不一致和先前迭代中特定领域知识的丢失。本文提出了一个在沃尔玛部署的统一管道,解决信号质量和模型演进问题。贡献有两方面:一是混合硬负样本挖掘,集成在线跨批次采样增加负样本多样性量级,结合离线挖掘识别细微不匹配;二是遗留感知蒸馏,从DistilBERT过渡到更高容量的GTE-base编码器,并引入热启动蒸馏技术转移特定领域专业知识。通过大量离线实验和在线A/B测试验证,该管道已投入生产,NDCG@5提高了7.34%,总收入提升了0.50%。
英文摘要
Embedding-based retrieval (EBR) is foundational to large-scale e-commerce search, yet its effectiveness is often constrained by the quality of training signals and the representational capacity of the encoder. Standard dual-encoders suffer from a training-inference gap: they are optimized on narrow candidate pools but must discriminate against hundreds of millions of items during inference. Furthermore, while transitioning to higher-capacity backbones can mitigate this gap, simply replacing a mature model can lead to inconsistent retrieval behavior and a loss of the domain-specific knowledge established in previous iterations. In this paper, we present a unified pipeline deployed at Walmart that addresses both signal quality and model evolution. Our contributions are two-fold: (1) Hybrid Hard Negative Mining: We integrate Online Cross-Batch Sampling to increase negative diversity by an order of magnitude and Hybrid Offline Mining, which combines cross-encoder predictions with metadata heuristics to identify nuanced mismatches. (2) Legacy-Aware Distillation: We transition from DistilBERT to a higher-capacity GTE-base encoder. To ensure a smooth and superior transition, we introduce a Warm-Start Distillation technique that transfers domain-specific expertise from the legacy model to the new backbone. Validated through extensive offline experiments and online A/B testing, the proposed pipeline is deployed in live production, delivering a +7.34% improvement in NDCG@5 and a +0.50% lift in gross revenue.