arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AutoResearch 在生产规模下的应用:失败模式与多智能体框架

AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework

Aparajith Chandran, Juwon Kim, Saurav Jha, Pablo Castells, Florian Hottier

arXiv 2609.30541首次发表:更新:

发表机构

Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文在生产规模下应用 AutoResearch 范式,通过多智能体框架解决五种失败模式,实现 1.82 倍 Recall@6 提升和 5.8 倍目录覆盖扩展,证明这些失败模式是结构性属性。

AI 中文摘要

优化生产推荐管道的嵌入系统需要系统性的探索,而这种探索在规模上消耗了不成比例的工程精力。我们将 Andrej Karpathy 的 AutoResearch 范式——即由大语言模型迭代编辑训练脚本,并保留能改进保留标量指标的修改——应用于自动化这一探索过程。我们报告了在生产规模下运行该范式十二周的结果,期间每次迭代消耗数小时的 GPU 多卡计算,评估涉及相互竞争的标准,且活动周期跨越数周、涉及众多训练任务。在面向图书推荐管道的两个独立开发的表示学习系统中,我们运行了 220 多个实验,并观察到了原始设定中不存在的五种反复出现的失败模式:基础设施脆弱性、智能体记忆衰减、搜索方向停滞、迭代成本不对称和指标固着。我们贡献了一个基于三原则的脚手架设计——预防、持久化、重定向——该设计将每种失败模式映射到结构性补救措施,其实例化随迭代成本扩展。该框架在手工调优基线上实现了 1.82 倍的 Recall@6 提升和 2.1 倍的连贯性提升,且智能体自主设计了一个纯文本回退方案,将目录覆盖范围扩大了 5.8 倍。这两个系统在单次迭代成本上相差近三个数量级,却表现出相同的失败模式,这表明这些是生产规模自主研究的结构性属性,而非任一应用的特定产物。

英文摘要

Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a training script and retains modifications that improve a held-out scalar metric -- to automate this exploration. We report on twelve weeks of running this paradigm at production scale, where iterations consume hours of multi-GPU compute, evaluation involves competing criteria, and campaigns span weeks across many training jobs. Across two independently developed representation-learning systems for a book recommendation pipeline, we ran 220+ experiments and observed five recurring failure modes absent from the original setting: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation. We contribute a three-principle scaffolding design -- prevent, persist, redirect -- that maps each failure mode to a structural remedy and whose instantiation scales with iteration cost. The framework produced a 1.82x Recall@6 lift and a 2.1x coherence lift over hand-tuned baselines, and the agent autonomously designed a text-only fallback that expanded catalog coverage by 5.8x. The two systems span nearly three orders of magnitude in per-iteration cost yet exhibit the same failure modes, suggesting these are structural properties of production-scale autonomous research rather than artifacts of either application.

Comments10 pages, 3 figures. Accepted at IEEE ICDM 2026 (Applied Track)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑