发表机构
Google(谷歌公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出RecEvolve知识驱动型自主智能体系统,用于优化双塔推荐模型,实现NDCG提升约20%、用户满意度提升3.77%,同时发现评估协议漏洞及新挑战。
AI 中文摘要
智能体AI的兴起推动了向自迭代系统的转变,为生产环境中推荐模型的自主优化开辟了新前沿。本文对一种知识驱动型自主智能体系统进行了实证验证,该系统直接部署在生产级大规模双塔(Two-Tower)检索模型上。通过将整个研究生命周期——从创意生成、代码实现、离线训练到指标评估——委托给一个连续闭环自主框架,该智能体系统从头完成了40余次自主训练运行。在严格的生产规模评估下,这些运行系统地解决了最新生产模型上的隐藏架构瓶颈,实现了NDCG指标约20%的相对提升,该增益直接转化为生产环境实时流量中用户满意度提升3.77%。此外,该部署还暴露了标准评估协议中的关键漏洞,因为智能体系统自主发现了奖励作弊捷径。这些发现证明,自主流水线可显著加速机器学习研究的步伐,并对基础实验基础设施的严谨性进行压力测试,同时也暴露了奖励作弊、对失败假设的冗余探索等新挑战。
英文摘要
The rise of agentic AI has catalyzed a shift toward self-iterating systems, opening new frontiers for the autonomous optimization of production recommender models. This paper presents the empirical validation of a knowledge-driven autonomous agent system, deployed directly on a production large-scale Two-Tower retrieval model. By delegating the entire research lifecycle, spanning idea generation, code implementation, offline training, and metric evaluation, to a continuous closed-loop autonomous framework, the agent system executed over 40 completed autonomous training runs from scratch. Executing these runs under rigorous production-scale evaluations, the system systematically navigated hidden architectural bottlenecks on the latest production model to achieve a breakthrough ~20% relative improvement in NDCG, a gain that translated directly to a +3.77% increase in user satisfaction in live production traffic. Furthermore, the deployment exposed critical vulnerabilities in standard evaluation protocols, as the agent system autonomously discovered reward-hacking shortcuts. These findings prove that an autonomous pipeline can dramatically accelerate the pace of machine learning research and stress-test the rigorousness of underlying experimental infrastructure, while also exposing novel challenges such as reward hacking and redundant exploration of failed hypotheses.
Comments8 pages, 4 figures, target conference: RecSys '26