arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SalesLoop:基于绩效反馈的强化学习用于销售线索排名

SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking

Chenyu Zhang

arXiv 2607.20655首次发表:更新:

发表机构

Li Auto Inc.(理想汽车公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对CRM系统中线索排名模型离线精度高但生产中表现不佳的问题,提出SalesLoop强化学习框架,通过引入性能感知奖励和判别式GRPO,有效提升了排名指标,经生产测试验证有显著效果。

AI 中文摘要

客户关系管理(CRM)系统中的线索排名面临持续挑战:离线精度高的模型在生产中往往表现不佳。我们识别出造成这种脱节的三个根本差距:离线-在线指标不匹配、逐点-逐列表目标不一致和时间分布漂移。为解决这些差距,我们提出SalesLoop,一种强化学习框架,在模型预测和实际业务结果之间建立封闭反馈回路。我们的方法引入了性能感知奖励和判别式GRPO。SalesLoop比最强的静态基线将NDCG@K提高了7.9%,P@K提高了15.8%。在一家新能源汽车制造商进行的160天生产A/B测试验证了统计学上显著的累计提升。在生产中,排名主干实现了44.1%的前10%召回率,并以专家基线转化率的2.3倍展示高意向线索。

英文摘要

Lead ranking in Customer Relationship Management (CRM) systems faces a persistent challenge: models achieving high offline accuracy often underperform in production. We identify three fundamental gaps responsible for this disconnect: offline-online metric mismatch, pointwise-listwise objective misalignment, and temporal distribution drift. To address these gaps, we propose SalesLoop, a reinforcement learning framework that establishes a closed feedback loop between model predictions and real-world business outcomes. Our approach introduces (1) a performance-aware reward that encodes conversion outcomes weighted by ranking position and conversion velocity, and (2) Discriminative GRPO, a listwise optimization objective that adapts Group Relative Policy Optimization to discriminative ranking models. SalesLoop improves NDCG@K by +7.9\% and P@K by +15.8\% over the strongest static baseline. A 160-day production A/B test at a New Energy Vehicle manufacturer, spanning 16.5M leads and 280 sales specialists across two provincial markets, validates statistically significant cumulative lift of +4.7\% ($p=0.047$) and +8.7\% ($p=0.002$). In production, the ranking backbone achieves Top-10\% recall of 44.1\% and surfaces high-intent leads at $2.3\times$ the conversion rate of specialist baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑