发表机构
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出LOFA框架,结合强化学习与反馈感知的策略内蒸馏,利用真实在线用户反馈改进购物智能体,在电商日志实验中提升了推荐质量、回复有用性及用户满意度对齐效果。
AI 中文摘要
基于大语言模型的购物智能体正越来越多地部署在真实电商平台中,产生的海量用户交互日志为改进这些智能体提供了宝贵的监督信号。然而,现有方法主要依赖离线训练信号,如用户-商品交互或合成偏好数据,却很大程度上忽略了用户自然对话反馈中蕴含的丰富监督信息。此外,可用的在线反馈具有异构性、稀疏性和噪声特性,难以自动转化为可靠的学习信号。为应对这些挑战,我们提出LOFA框架,该框架无需人工标注即可让购物智能体直接从真实在线交互日志中学习。LOFA将基于可验证购买结果的强化学习与感知反馈的策略内蒸馏相结合,后者可识别用户对话内指令并将其转化为密集的词元级监督信号。这些互补目标同时捕捉了协同行为模式和用户特定偏好。在真实电商日志上开展的大量实验表明,与强基线方法相比,LOFA在推荐质量、回复有用性和用户满意度对齐方面均实现了持续提升,凸显了从真实在线用户反馈中学习购物智能体的有效性。
英文摘要
Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating massive amounts of user interaction logs that provide valuable supervision for improving these agents. However, existing approaches primarily rely on offline training signals, such as user-item interactions or synthetic preference data, while largely overlooking the rich supervision contained in users' natural conversational feedback. Moreover, the available online feedback is heterogeneous, sparse, and noisy, making it difficult to transform into reliable learning signals automatically. To address these challenges, we propose LOFA, a framework that enables shopping agents to learn directly from real online interaction logs without human annotation. LOFA combines reinforcement learning over verifiable purchase outcomes with feedback-aware on-policy distillation, which identifies users'in-dialogue directives and converts them into dense token-level supervision. These complementary objectives capture both collaborative behavioral patterns and user-specific preferences. Extensive experiments on real-world e-commerce logs demonstrate that LOFA consistently improves recommendation quality, response helpfulness, and user-satisfaction alignment over strong baselines, highlighting the effectiveness of learning shopping agents from real online user feedback.