arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PILOT技术报告

PILOT Technical Report

Jiuning Lin, Ruiquan Lan, Xiaodong Zhu, Bin Zhang, Chengyu Lai, Chuxin Chen, Dimin Wang, Han Zhu, Hongtao Cheng, Jialin Zhu, Lingqing Zhang, Shuai Zhong, Tao Wang, Weipeng Huang, Yinjiang Cai, Yinnan Song, Yuan Liu, Zhibo Xiao, Zhixin Ma, Zihong Huang

arXiv 2608.18637首次发表:更新:

AI 中文总结

该研究针对现有推荐系统优化智能体方法的反应式缺陷,提出PILOT框架,通过三类角色构建控制循环,在淘宝平台对比ROAM验证,实现多项指标提升且搜索效率显著提高,无需人工干预。

AI 中文摘要

现有的用于推荐系统优化的智能体方法本质上仍是反应式的:它们会根据观测到的指标变化调整参数,但缺乏主动设计受控实验、在用户细分层面个性化策略,或在不同任务间积累可复用实验方法的能力。我们提出PILOT(Proactive Insight Learner for Online Tree-Experiments,在线树实验主动洞察学习器),这是一个大语言模型智能体框架,在受约束的控制循环内组织三个角色,其中确定性服务强制执行所有安全、统计和权限边界:(1)实验管理器,仅从规则生成的合法命令包中选择内容,驱动完整实验生命周期——任务接收、观测治理、异常恢复和事后分析;(2)搜索规划器,仅在管理器请求规划时被调用,提出用于用户细分层面个性化的候选决策树;(3)记忆策展人,异步将实验结果提炼为策略层面的领域知识和带有溯源追踪的方法,与主循环隔离以实现故障隔离。管理器使智能体具备主动性,规划器支持超越全局调优的群体层面个性化,策展人将每个完成的任务转化为下一个任务的学习机会。PILOT在淘宝平台部署,包含5个实验桶,与ROAM(Reactive Optimization with Agent-driven Moves,基于智能体驱动动作的反应式优化,一种无生命周期治理或结构化假设检验的自由探索智能体)对比。PILOT实现了最高+1.40%的IPV、+1.60%的核心IPV、+0.96%的交易数量和+1.50%的交易金额,优于ROAM的最佳结果(+1.00%的IPV、+0.90%的核心IPV、+0.60%的交易数量、+1.13%的交易金额),同时将搜索效率从53.3%提升至93.3%(提升40个百分点),整个实验周期无需人工干预。

英文摘要

Existing agentic approaches for recommendation system optimization remain fundamentally reactive: they adjust parameters in response to observed metric changes but lack the ability to proactively design controlled experiments, personalize strategies at the user-segment level, or accumulate reusable experimental methodology across tasks. We present PILOT (Proactive Insight Learner for Online Tree-Experiments), an LLM-agent framework that organizes three roles within a constrained control loop where deterministic services enforce all safety, statistical, and permission boundaries: (1) an Experiment Manager that drives the full experiment lifecycle -- task intake, observation governance, anomaly recovery, and postmortem -- by selecting only from a rule-generated legal-command envelope; (2) a Search Planner that proposes candidate decision trees for user-segment-level personalization, invoked only when the Manager requests planning; and (3) a Memory Curator that asynchronously distills experiment outcomes into strategy-level domain knowledge and provenance-tracked methodology, failure-isolated from the main loop. The Manager makes the agent proactive, the Planner enables population-level personalization beyond global tuning, and the Curator turns every completed task into a learning opportunity for the next. Deployed on Taobao's platform with 5 experimental buckets, PILOT is compared against ROAM(Reactive Optimization with Agent-driven Moves), a free-exploration agent without lifecycle governance or structured hypothesis testing. PILOT achieves up to +1.40% IPV, +1.60% Core IPV, +0.96% transaction count, and +1.50% transaction amount, improving over ROAM's best results (+1.00% IPV, +0.90% Core IPV, +0.60% transaction count, +1.13% transaction amount) while raising search efficiency from 53.3% to 93.3% (+40 pp), with no human intervention throughout the experimental cycle.

CommentsTechnical Report, 42 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑