发表机构
The Hong Kong Polytechnic University; Kuaishou Technology; University of Electronic Science and Technology of China; Southwest Jiaotong University(香港理工大学; 快手科技; 电子科技大学; 西南交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出A/B Agent智能体,通过层级经验树与Tree-RAG技术优化工业A/B测试的策略迭代,在短视频电商场景实现GMV提升4.829%且护栏指标正向增益。
AI 中文摘要
工业推荐策略迭代高度依赖大规模A/B实验。传统调优需要专家反复设计策略、配置实验、分析结果并调整参数,过程费力且耗时。同时,历史实验中的宝贵知识往往碎片化,仅靠人工专家难以实现系统复用。现有RAG智能体通过检索过往策略部分减轻了这一负担,但通常以扁平方式组织经验,忽略业务场景、推荐阶段、优化目标和实验上下文之间的层级关系,这常导致检索不匹配和跨场景迁移受限,还阻碍智能体通过连续A/B反馈持续优化策略和参数。为解决这些局限,我们提出A/B Agent,一种用于工业推荐策略优化的闭环A/B智能体。该框架包含三个紧密耦合的核心组件:历史策略知识组织、自主目标感知策略生成、实验引导策略自进化。它将历史策略组织为层级经验树,通过多路径Tree-RAG检索可迁移证据以生成可执行策略,并持续分析在线A/B反馈以指导自主调优、更新经验树实现自进化。大量离线和在线评估证明了其有效性,包括在真实短视频电商推荐系统中GMV提升4.829%,同时所有护栏指标均保持正向增益。
英文摘要
Industrial recommendation strategy iteration heavily relies on large-scale A/B experimentation. Traditional tuning requires experts to repeatedly design strategies, configure experiments, analyze results, and adjust parameters, making the process labor-intensive and time-consuming. Meanwhile, valuable knowledge from historical experiments is often fragmented, making systematic reuse difficult through manual expert effort alone. Existing RAG agents partially alleviate this burden by retrieving prior strategies, but typically organize experience in a flat manner, overlooking the hierarchical relationships among business scenarios, recommendation stages, optimization objectives, and experimental contexts. This often results in mismatched retrieval and limited cross-scenario transfer, while preventing agents from continuously refining strategies and parameters through sequential A/B feedback. % To address these limitations, we propose A/B Agent, a closed-loop A/B agent for industrial recommendation strategy optimization. The framework comprises three tightly coupled core components: Historical Strategy Knowledge Organization, Autonomous Target-Aware Strategy Generation, and Experiment-Guided Strategy Self-Evolution. It organizes historical strategies into a hierarchical experience tree, retrieves transferable evidence through multi-path Tree-RAG to generate executable strategies, and continuously analyzes online A/B feedback to guide autonomous tuning and update the experience tree for self-evolution. Extensive offline and online evaluations demonstrate its effectiveness, including a 4.829% improvement in GMV in a real-world short-video e-commerce recommendation system while maintaining positive gains across all guardrail metrics.