推进AgentX中的模型研究:面向工业推荐系统的长时程自主性
Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems
浏览论文内容
中文总结 AI 辅助
AgentX-Model通过双智能体架构和四个研究动作,在工业推荐系统中实现长时程自主模型研究,生产评估中560/636实验AUC超基线,在线A/B测试带来多项业务收益。
中文摘要 AI 辅助
维持工业推荐研究需要利用一个实验的结果来决定下一步研究什么。我们提出了AgentX-Model,即AgentX模型研究框架的下一代,它在由业务输入和预测任务定义的沙盒内连接提案开发与模型实验。AgentX-Model采用双智能体架构,包含一个研究智能体和一个模型智能体。研究智能体根据论文和实验发现制定经过独立评审的提案,而模型智能体进行多轮调查并返回代码、测量结果和未解决的问题。利用返回的结果,研究智能体选择一个起始实现并制定下一个研究问题,使后续实验能够基于先前发现进行构建。我们将这一持续研究组织为四个动作:复现、跟进、组合和诊断。前三个动作驱动常规研究,而诊断获取选择修复所需的证据,包括针对业务反馈和在线评估提出的问题,如通过PCOC测量的预测偏差。在生产评估中,636个完成模型变更的实验中有560个记录的AUC高于其业务基线。随着研究的继续,一些实验记录的AUC高于其谱系中所有可比较的祖先。最近在不同业务设置中进行的五次在线A/B评估报告了收益,包括获取效率提升10-15%,目标细分广告支出提升15-20%,观看时长提升0.3-0.8%;观看时长模型使用了约10%更少的FLOPs和参数。一个依赖感知的历史回放基准进一步评估研究分配,初步结果显示,当智能体已经分析和选择具体候选时,更复杂的调度并未带来一致的效率提升。
英文摘要
Sustaining industrial recommendation research requires using the results of one experiment to decide what to investigate next. We present AgentX-Model, the next generation of AgentX's model research framework, which connects proposal development and model experimentation within sandboxes defined by business inputs and prediction tasks. AgentX-Model adopts a dual-agent architecture comprising a Research Agent and a Model Agent. The Research Agent develops independently reviewed proposals from papers and experimental findings, while the Model Agent conducts multi-round investigations and returns code, measurements, and unresolved questions. Using the returned results, the Research Agent selects a starting implementation and formulates the next research question, allowing subsequent experiments to build on earlier findings. We organize this continuing research around four actions: Reproduce, Follow-up, Composition, and Diagnose. The first three actions drive routine research, while Diagnose acquires the evidence needed to choose a repair, including for issues raised by business feedback and online evaluation, such as prediction bias measured by PCOC. Across the production evaluation, 560 of 636 completed model-changing experiments recorded AUC above their business baselines. As research continued, some experiments recorded AUC above every comparable ancestor in their lineages. The five latest online A/B evaluations across different business settings reported gains including 10-15% in acquisition efficiency, 15-20% in target-segment advertising spend, and 0.3-0.8% in watch time; the watch-time model used approximately 10% fewer FLOPs and parameters. A dependency-aware historical-replay benchmark further evaluates research allocation, with initial results showing no consistent efficiency gain from more complex scheduling when agents already analyze and select concrete candidates.