发表机构
Meta Platforms, Inc.(元平台公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对广告排序中人类ML迭代瓶颈,提出A-MLE自主LLM代理系统,分解迭代为五阶段并含人工检查点,实验表明其是工程师的实用倍增器。
AI 中文摘要
现代工业广告排序栈的瓶颈日益不再源于模型容量或训练计算,而是源于人类机器学习迭代的吞吐量——即为了获得一个统计上显著的改进所需的研究、实现、训练、调试、评估和上线循环。典型的排序栈包含众多异构数据、架构和基础设施约束的差异化模型,每个循环每个模型需要资深工程师数天到数周的时间。因此,在某一模型上被证明有效的技术向其他模型的扩散缓慢且不均衡,留下了大量可恢复的信号未被探索。我们提出了 Agentic ML Exploration (A-MLE),一个自主的 LLM 代理系统,系统地探索跨广告排序模型组合的机器学习技术。A-MLE 将机器学习迭代分解为五个阶段,涉及假设生成、探索策略、实验执行、结果分析和共享知识基础,这些阶段由一个代理编排,该代理在沙箱执行层上调用领域特定技能和代理工作流,并在每个阶段边界设置人工介入检查点。我们在代表性的大规模广告排序模型集上部署了 A-MLE,并沿分层能力框架(工具可用性、自主工作流执行和开放式探索)对其进行评估。我们进一步报告了一项使用固定代理循环的受控跨 LLM 研究,该研究揭示了 Claude Sonnet、Gemini 和 GPT 系列在执行可靠性和探索激进性方面的定性差异。我们讨论了失败模式和决定可靠性的设计选择。我们的研究结果表明,代理探索是工业推荐系统中机器学习工程师的实用倍增器,特别是对于很少获得专家关注的尾部模型。
英文摘要
Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute, but by the throughput of human ML iteration - the cycles of research, implementation, training, debugging, evaluation, and launch required to surface a single statistically significant improvement. A typical ranking stack contains numerous differentiated models with heterogeneous data, architectures, and infrastructure constraints, and each cycle takes days to weeks of senior engineer attention per model. As a result, techniques that have proven effective on one model diffuse into others slowly and unevenly, leaving substantial recoverable signal unexplored. We present Agentic ML Exploration (A-MLE), an autonomous LLM-agent system that systematically explores ML techniques across a portfolio of ads ranking models. A-MLE decomposes ML iteration into five stages involving hypothesis generation, exploration strategy, experiment execution, result analysis and shared knowledge substrate which are orchestrated by a single agent that invokes domain-specific skills and agentic workflows against a sandboxed execution layer, with human-in-the-loop checkpoints at each stage boundary. We deploy A-MLE across a representative set of large-scale ads ranking models and evaluate it along a tiered capability framework (tool availability, autonomous workflow execution, and open-ended exploration). We further report a controlled cross-LLM study using a fixed agent loop, which surfaces qualitative differences in execution reliability and exploration aggressiveness across the Claude Sonnet, Gemini, and GPT families. We discuss failure modes and the design choices that govern reliability. Our findings suggest that agentic exploration is a practical force multiplier for ML engineers in industrial recommenders, especially for the long tail of models that rarely receive expert attention.
Comments7 pages, 4 figures