AI 中文总结
该研究针对现有推荐系统优化智能体方法的反应式缺陷,提出PILOT框架,通过三类角色构建控制循环,在淘宝平台对比ROAM验证,实现多项指标提升且搜索效率显著提高,无需人工干预。
AI 中文摘要
现有的用于推荐系统优化的智能体方法本质上仍是反应式的:它们会根据观测到的指标变化调整参数,但缺乏主动设计受控实验、在用户细分层面个性化策略,或在不同任务间积累可复用实验方法的能力。我们提出PILOT(Proactive Insight Learner for Online Tree-Experiments,在线树实验主动洞察学习器),这是一个大语言模型智能体框架,在受约束的控制循环内组织三个角色,其中确定性服务强制执行所有安全、统计和权限边界:(1)实验管理器,仅从规则生成的合法命令包中选择内容,驱动完整实验生命周期——任务接收、观测治理、异常恢复和事后分析;(2)搜索规划器,仅在管理器请求规划时被调用,提出用于用户细分层面个性化的候选决策树;(3)记忆策展人,异步将实验结果提炼为策略层面的领域知识和带有溯源追踪的方法,与主循环隔离以实现故障隔离。管理器使智能体具备主动性,规划器支持超越全局调优的群体层面个性化,策展人将每个完成的任务转化为下一个任务的学习机会。PILOT在淘宝平台部署,包含5个实验桶,与ROAM(Reactive Optimization with Agent-driven Moves,基于智能体驱动动作的反应式优化,一种无生命周期治理或结构化假设检验的自由探索智能体)对比。PILOT实现了最高+1.40%的IPV、+1.60%的核心IPV、+0.96%的交易数量和+1.50%的交易金额,优于ROAM的最佳结果(+1.00%的IPV、+0.90%的核心IPV、+0.60%的交易数量、+1.13%的交易金额),同时将搜索效率从53.3%提升至93.3%(提升40个百分点),整个实验周期无需人工干预。
英文摘要
Existing agentic approaches for recommendation system optimization remain fundamentally reactive: they adjust parameters in response to observed metric changes but lack the ability to proactively design controlled experiments, personalize strategies at the user-segment level, or accumulate reusable experimental methodology across tasks. We present PILOT (Proactive Insight Learner for Online Tree-Experiments), an LLM-agent framework that organizes three roles within a constrained control loop where deterministic services enforce all safety, statistical, and permission boundaries: (1) an Experiment Manager that drives the full experiment lifecycle -- task intake, observation governance, anomaly recovery, and postmortem -- by selecting only from a rule-generated legal-command envelope; (2) a Search Planner that proposes candidate decision trees for user-segment-level personalization, invoked only when the Manager requests planning; and (3) a Memory Curator that asynchronously distills experiment outcomes into strategy-level domain knowledge and provenance-tracked methodology, failure-isolated from the main loop. The Manager makes the agent proactive, the Planner enables population-level personalization beyond global tuning, and the Curator turns every completed task into a learning opportunity for the next. Deployed on Taobao's platform with 5 experimental buckets, PILOT is compared against ROAM(Reactive Optimization with Agent-driven Moves), a free-exploration agent without lifecycle governance or structured hypothesis testing. PILOT achieves up to +1.40% IPV, +1.60% Core IPV, +0.96% transaction count, and +1.50% transaction amount, improving over ROAM's best results (+1.00% IPV, +0.90% Core IPV, +0.60% transaction count, +1.13% transaction amount) while raising search efficiency from 53.3% to 93.3% (+40 pp), with no human intervention throughout the experimental cycle.
CommentsTechnical Report, 42 pages, 10 figures