arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SR-Agent:一种用于电子商务推荐中排序后策略优化的经验驱动智能框架

SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategy Refinement in E-Commerce Recommendation

Hanchen Yang, Kaiwen Yang, Junpeng Zhuang, Yang He, Keting Cen, Bochao Liu, Zhongbo Sun, An Liu, Zhongteng Han, Chenyi Lei

arXiv 2607.17719首次发表:更新:

发表机构

Kuaishou Technology(快手科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

电商推荐系统中,后排序策略因环境变化需优化,以往方式存在问题。本文提出SR-Agent框架,统一用户模拟、分析和策略优化组件,经快手平台测试,能提升订单量、浏览深度和点击类别多样性,还缩短优化周期、降低成本。

AI 中文摘要

用户体验是工业电子商务推荐系统中的首要目标。后排序策略因其简单性和低服务成本而广泛应用于工业推荐系统中,但随着在线推荐环境的不断发展,这些静态配置的策略逐渐过时,降低了用户体验。改进这些策略通常依赖于人工检查、诊断和更新,这一过程缓慢、成本高且难以重用。尽管最近基于大语言模型的智能体提供了有前景的方向,但它们都没有实现自动化、自我进化的策略改进的完整闭环。为了弥补这一差距,我们引入了SR-Agent,这是一个策略改进智能体框架,据我们所知,它是第一个用于工业推荐系统中改进后排序策略的框架。SR-Agent统一了三个组件:(i)一个用户模拟智能体,它应用分阶段检查技能来发现用户感知到的不良案例;(ii)一个分析智能体,它将反复出现的不良案例整合为结构化、可重用的诊断;(iii)一个受约束的策略改进工具,它将诊断映射到类型化和有界的行动,由一个具有可逆回滚的四阶段奖励管道控制。在快手电子商务平台上部署后,SR-Agent持续运行此改进循环,在为期一个月的在线A/B测试中,订单量增加了0.71%,浏览深度增加了0.34%,点击类别多样性增加了0.48%,同时显著缩短了改进周期并降低了运营成本。

英文摘要

User experience is a first-class objective in industrial e-commerce recommender systems (RS). Post-ranking strategies, which govern diversity, similarity, and exposure over a ranked list, are widely deployed in industrial RS for their simplicity and low serving cost. However, as the online recommendation environment evolves continuously, these statically configured strategies gradually become stale, thereby degrading the user experience. Refining them typically relies on manual inspection, diagnosis, and updates, making it slow, costly, and difficult to scale or reuse. Although recent LLM-based agents (e.g., RecUserSim, SimUSER, and Self-EvolveRec) offer promising directions, none of them close the full loop of automated, self-evolving strategy refinement. To bridge this gap, we introduce SR-Agent, which, to the best of our knowledge, is the first agentic framework deployed to refine post-ranking strategies in industrial RS. SR-Agent unifies three components: (i) a UserSim agent that applies inspection skills to surface user-perceived bad cases; (ii) an Analysis agent that consolidates recurring bad cases into structured, reusable diagnoses; and (iii) a constrained Strategy Refinement Harness that maps diagnoses to typed and bounded actions, gated by a four-stage reward pipeline with reversible rollback. Deployed on the Kuaishou e-commerce platform, SR-Agent continuously runs this refinement loop and, in a one-month online A/B test, increases order volume by 0.71%, browsing depth by 0.34%, and clicked-category diversity by 0.48%, while markedly shortening the refinement cycle and lowering operational cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑