arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26921cs.IR

将词汇产品关联提炼到深度 Transformer 中:面向自然语言电子商务搜索的极端多标签方法

Distilling Lexical Product Associations into Deep Transformers: An Extreme Multi-Label Approach for Natural Language E-Commerce Search

  • International Institute of Information Technology Bangalore(班加罗尔国际信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

Sunnidhya Roy, Samarpita Bhaumik

AI总结:

针对电子商务搜索中词汇不匹配问题,提出将词汇产品关联蒸馏到 DistilBERT 的极端多标签方法,在 54,000 产品上实现高精度,优于词汇模型并支持工业规模部署。

AI中文摘要:

传统的电子商务搜索平台严重依赖倒排索引和基于词元的词汇匹配算法(如 BM25 和 TF-IDF),这些方法在处理对话式、意图驱动或改写后的用户查询时经常失效——这是经典的词汇不匹配问题。我们将对话式产品推荐表述为一个极端多标签分类(XMLC)问题,其目录来自 Amazon Reviews '23 基准,包含 N = 54,000 个产品,涵盖 27 个平衡的零售类别。我们使用预训练的 DistilBERT Transformer 编码器,通过伪标签知识蒸馏框架,将密集的物品间相似性拓扑(基于累积元数据通过 TF-IDF 余弦相似度生成,取 K = 50 个最近邻)蒸馏到深度上下文表示中。在严格的 85/15 训练/验证划分(8,089 个保留产品,C = 53,923 个输出类别)上评估,并强制执行严格的自我排除,DistilBERT 神经学生模型实现了 P@1 = 93.15%,P@5 = 90.08%,NDCG@10 = 0.8845,MRR@10 = 0.9545,紧密接近由修正后的 TF-IDF 教师模型建立的经验上限(P@1 = 98.10%,NDCG@10 = 0.9419,MRR@10 = 0.9882)。此外,针对十种结构化自然语言查询原型(涵盖情境性、跨类别、改写和否定约束查询)的定性基准表明,Transformer 学生模型在关键词匹配之外实现了显著泛化,成功解决了词汇模型完全失败的隐式用户意图。最后,我们分析了在工业目录规模(> 10^6 个物品)下极端分类投影层的架构和内存可扩展性权衡,并提出了迈向双编码器(双塔)向量搜索的具体部署路径。代码:https://github.com/Sunnidhya/Distilling-Lexical-Product-Associations-into-Deep-Transformers。

英文摘要:

Traditional e-commerce search platforms rely heavily on inverted indices and token-level lexical matching algorithms (e.g., BM25 and TF-IDF), which frequently fail on conversational, intent-driven, or paraphrased user queries -- the classic vocabulary mismatch problem. We formulate conversational product recommendation as an Extreme Multi-Label Classification (XMLC) problem over an e-commerce catalog of N = 54,000 products spanning 27 balanced retail categories from the Amazon Reviews '23 benchmark. Using a pre-trained DistilBERT transformer encoder, we distill dense item-to-item similarity topologies (generated via TF-IDF cosine similarity over cumulative metadata with K = 50 nearest neighbours) into a deep contextual representation via a pseudo-label knowledge distillation framework. Evaluated on an exact 85/15 train/validation split (8,089 held-out products across C = 53,923 output classes) with strict self-exclusion enforced, the DistilBERT neural student achieves P@1 = 93.15%, P@5 = 90.08%, NDCG@10 = 0.8845, and MRR@10 = 0.9545, closely recovering the empirical ceiling established by the corrected TF-IDF teacher (P@1 = 98.10%, NDCG@10 = 0.9419, MRR@10 = 0.9882). Furthermore, a qualitative benchmark across ten structured natural language query archetypes -- encompassing situational, cross-category, paraphrased, and negative-constraint queries -- demonstrates that the transformer student generalises substantially beyond keyword matching, successfully resolving implicit user intent where lexical models fail completely. Finally, we analyse the architectural and memory scalability trade-offs of extreme classification projection layers at industrial catalog scale (> 10^6 items) and present a concrete deployment trajectory toward Dual-Encoder (Two-Tower) vector search. Code: https://github.com/Sunnidhya/Distilling-Lexical-Product-Associations-into-Deep-Transformers.

补充信息

↑