通过LLM蒸馏进行多标签主题分配:生成式与判别式学生模型的比较分析
Multi-Label Topic Assignment via LLM Distillation: A Comparative Analysis of Generative vs. Discriminative Student Models
浏览论文内容
中文总结 AI 辅助
本研究通过比较生成式与判别式学生模型在多标签主题分配中的性能,发现生成式模型在复杂对话和标签扩展场景下更优,并已成功部署于大规模电商生产环境。
中文摘要 AI 辅助
针对用户生成内容(UGC)的多标签主题分配——包括产品评论和买卖双方对话——由于非正式语言、极端标签稀疏性以及快速演变的分类体系,在大规模电子商务中带来了独特的可扩展性挑战。虽然利用大型语言模型(LLMs)作为标注预言机来蒸馏真实数据已成为行业标准,以绕过高昂的人工标注成本,但确定结果学生模型的最优低延迟架构仍然是一个未解决的挑战。为了解决这一问题,我们在小语言模型(SLM)参数规模(1B、4B和8B)和架构范式(因果生成式与双向判别式)上进行了全面评估。通过将生成式文本到标签分类器与判别式基线(DeBERTa-V3和ModernBERT)进行比较,我们的分析揭示了一个关键的数据依赖权衡:虽然判别式模型在结构化产品评论上优于超轻量级生成式模型,但即使是最小的1B生成式模型在复杂的多轮对话数据上也超过了判别式基线。此外,生成式模型在大量标签集扩展(最多112个主题)和严重的长尾分布下保持了稳健性能,而判别式基线在规模扩大时Macro-F1下降了35%。最后,我们详细介绍了这些优化模型在产品评论和对话领域的成功生产部署,展示了在全球市场范围内的严格延迟合规性和切实的业务影响。
英文摘要
Multi-label topic assignment for user-generated content (UGC) -- including product reviews and buyer-seller conversations -- poses unique scalability challenges in large-scale e-commerce due to informal language, extreme label sparsity, and rapidly evolving taxonomies. While utilizing Large Language Models (LLMs) as labeling oracles to distill ground-truth data has emerged as an industry standard to bypass prohibitive manual annotation costs, determining the optimal, low-latency architecture for the resulting student models remains an open challenge. To address this, we conduct a comprehensive evaluation across Small Language Model (SLM) parameter scales (1B, 4B, and 8B) and architectural paradigms (causal generative versus bidirectional discriminative). Comparing generative text-to-label classifiers against discriminative baselines (DeBERTa-V3 and ModernBERT), our analysis reveals a crucial data-dependent trade-off: while discriminative models outperform ultra-lightweight generative models on structured product reviews, even the smallest 1B generative model surpasses discriminative baselines on complex, multi-turn conversational data. Furthermore, generative models maintain robust performance under massive label-set expansion (up to 112 topics) and severe long-tail distributions, whereas discriminative baselines suffer a 35% drop in Macro-F1 at scale. Finally, we detail the successful production deployment of these optimized models across both product review and conversational domains, demonstrating strict latency compliance and tangible business impact at a global marketplace scale.
发表机构
- eBay Inc.(eBay公司)
机构由 AI 辅助整理,请以论文原文为准。