arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30001cs.AIcs.IR

推进AgentX中的模型研究:面向工业推荐系统的长时程自主性

Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems

Shuang Yang, Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Yusheng Huang, Han Gao, Guanchen Wang, Tianbao Ma, Linxun Chen, Peilin Song, Xuming Wang, Chen Li… 展开作者

Shuang Yang, Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Yusheng Huang, Han Gao, Guanchen Wang, Tianbao Ma, Linxun Chen, Peilin Song, Xuming Wang, Chen Li, Fan Wu, Tao Wang, Zibo Zhao, Xiangyu Wu, An Liu, Fei Pan, Peng Jiang, Chen Yang, Zhaojie Liu, Wenwu Ou

首次发表
浏览论文内容

中文总结 AI 辅助

AgentX-Model通过双智能体架构和四个研究动作,在工业推荐系统中实现长时程自主模型研究,生产评估中560/636实验AUC超基线,在线A/B测试带来多项业务收益。

中文摘要 AI 辅助

维持工业推荐研究需要利用一个实验的结果来决定下一步研究什么。我们提出了AgentX-Model,即AgentX模型研究框架的下一代,它在由业务输入和预测任务定义的沙盒内连接提案开发与模型实验。AgentX-Model采用双智能体架构,包含一个研究智能体和一个模型智能体。研究智能体根据论文和实验发现制定经过独立评审的提案,而模型智能体进行多轮调查并返回代码、测量结果和未解决的问题。利用返回的结果,研究智能体选择一个起始实现并制定下一个研究问题,使后续实验能够基于先前发现进行构建。我们将这一持续研究组织为四个动作:复现、跟进、组合和诊断。前三个动作驱动常规研究,而诊断获取选择修复所需的证据,包括针对业务反馈和在线评估提出的问题,如通过PCOC测量的预测偏差。在生产评估中,636个完成模型变更的实验中有560个记录的AUC高于其业务基线。随着研究的继续,一些实验记录的AUC高于其谱系中所有可比较的祖先。最近在不同业务设置中进行的五次在线A/B评估报告了收益,包括获取效率提升10-15%,目标细分广告支出提升15-20%,观看时长提升0.3-0.8%;观看时长模型使用了约10%更少的FLOPs和参数。一个依赖感知的历史回放基准进一步评估研究分配,初步结果显示,当智能体已经分析和选择具体候选时,更复杂的调度并未带来一致的效率提升。

英文摘要

Sustaining industrial recommendation research requires using the results of one experiment to decide what to investigate next. We present AgentX-Model, the next generation of AgentX's model research framework, which connects proposal development and model experimentation within sandboxes defined by business inputs and prediction tasks. AgentX-Model adopts a dual-agent architecture comprising a Research Agent and a Model Agent. The Research Agent develops independently reviewed proposals from papers and experimental findings, while the Model Agent conducts multi-round investigations and returns code, measurements, and unresolved questions. Using the returned results, the Research Agent selects a starting implementation and formulates the next research question, allowing subsequent experiments to build on earlier findings. We organize this continuing research around four actions: Reproduce, Follow-up, Composition, and Diagnose. The first three actions drive routine research, while Diagnose acquires the evidence needed to choose a repair, including for issues raised by business feedback and online evaluation, such as prediction bias measured by PCOC. Across the production evaluation, 560 of 636 completed model-changing experiments recorded AUC above their business baselines. As research continued, some experiments recorded AUC above every comparable ancestor in their lineages. The five latest online A/B evaluations across different business settings reported gains including 10-15% in acquisition efficiency, 15-20% in target-segment advertising spend, and 0.3-0.8% in watch time; the watch-time model used approximately 10% fewer FLOPs and parameters. A dependency-aware historical-replay benchmark further evaluates research allocation, with initial results showing no consistent efficiency gain from more complex scheduling when agents already analyze and select concrete candidates.

补充信息

↑