LLM4AIGQ:面向多兴趣挖掘的基于大语言模型的AI引导查询生成框架
LLM4AIGQ: LLM-based AI Guidance Query Generation Framework for Multi Interest Mining
浏览论文内容
中文总结 AI 辅助
针对传统AIGQ生成方法存在的语义漂移、多兴趣挖掘不足等问题,提出基于LLM的LLM4AIGQ框架,采用SFT、RL、DPO训练及多级奖励设计,经离线评估和在线A/B测试验证了其性能。
中文摘要 AI 辅助
引导查询通过提取用户偏好来提供具有引导价值的搜索查询,从而刺激用户消费,在电子商务领域发挥着关键作用。传统AI生成查询(AIGQ)的生成主要依赖两阶段的“查询到AI生成查询(Q2AIGQ)”关联范式:首先通过多路径检索从用户画像、历史行为序列、商品侧信息及当前查询中召回用户的主要搜索查询,再通过基于规则的方法生成AIGQ。该方法因信息级联损失存在语义漂移问题;此外,主要搜索查询的推导严重依赖“用户-商品”共现关系,缺乏对用户多兴趣的挖掘,导致引导查询价值低且与购买意图不匹配。为解决传统基于共现的检索的表达局限性,我们提出LLM4AIGQ,一种针对用户多兴趣生成AI引导查询的基于大语言模型的解决方案。该方法通过整合用户画像和历史交互序列对用户兴趣进行分割,为每个子兴趣推断具体的消费意图,随后生成对应的AIGQ。在模型训练方面,我们采用包含监督微调(SFT)、强化学习(RL)和直接偏好优化(DPO)的后训练流程,以增强模型生成AIGQ的能力;还引入了多级奖励设计,以满足实际应用中多目标优化和长链推理的需求。在部署方面,我们采用近线生成、在线读取的架构,以满足延迟约束。大量实验分析表明,我们的模型在离线评估和在线A/B测试中均实现了稳健的性能。
英文摘要
Guidance queries stimulate user consumption by extracting preferences to provide search queries with guidance value, playing a crucial role in the e-commerce field. Traditional AI-generated queries (AIGQ) generation primarily relies on a two-stage "Query-to-AI-Generated-Query" (Q2AIGQ) association paradigm, first recalling user primary search queries from user profiles, historical behavior sequences, item-side information, and the current query through multi-path retrieval, then generalizing AIGQ via rule-based methods. This approach suffers from semantic drift due to information cascade loss; additionally, primary search query derivation heavily depends on "user-item" co-occurrence relationships, lacking exploration of user multi-interests, resulting in guidance queries with low value and mismatched purchase intent. To address the expressive limitations of traditional co-occurrence-based retrieval, we propose LLM4AIGQ, an LLM-based solution for generating AI guidance queries tailored to users' multi-interests. This approach segments user interests by integrating user profiles and historical interaction sequences, infers specific consumption intents for each sub-interest, and subsequently generates corresponding AIGQ. In terms of model training, we employ a post-training pipeline comprising Supervised Fine-Tuning (SFT), Reinforcement Learning (RL), and Direct Preference Optimization (DPO) to enhance the model's capability in generating AIGQ. We also introduce a multi-level reward design to satisfy the requirements of multi-objective optimization and long-chain reasoning in practical applications. Regarding deployment, we adopt a nearline-generation and online-read architecture to meet latency constraints. Extensive experimental analyses demonstrate that our model achieves robust performance in both offline evaluations and online A/B tests.
发表机构
- Huazhong University of Science and Technology(华中科技大学)
- Alibaba Group(阿里巴巴集团)
- Nankai University(南开大学)
机构由 AI 辅助整理,请以论文原文为准。