发表机构
Meta; University of Illinois Urbana-Champaign(Meta; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Auto-RecSys利用自主研究智能体,通过分布式异步执行、集中式内存和认知-程序分离,解决工业级推荐系统长反馈循环与系统复杂性挑战,显著减少人工时间并提升执行可靠性。
AI 中文摘要
自动研究智能体已展现出在假设生成、实验执行和迭代优化方面实现自动化的潜力。然而,将这一范式扩展到工业级推荐模型面临两大挑战:(1)长反馈循环,模型训练可能需要数天时间,使得串行迭代慢得难以接受,因此需要跨多个研究方向进行并行探索;(2)系统复杂性,大规模配置、脆弱的基础设施依赖以及多天GPU作业要求稳健且可恢复的执行。我们提出了Auto-RecSys,一个面向工业级推荐模型的长时域实验自主研究系统。Auto-RecSys通过三种框架设计应对这些挑战:(1)分布式异步执行,在多个服务器上并行运行多个实验;(2)集中式跨服务器内存,用于跨会话和故障的持久且可恢复的执行;(3)认知-程序分离,其中自然语言技能文件指导大语言模型推理,而确定性脚本强制执行操作正确性。Auto-RecSys进一步采用双循环自进化架构:执行进化循环,其中特定于模型的剧本通过记录失败尝试和固化成功流程来积累操作知识;以及想法进化循环,其中实验结果指导后续构思。在推荐模型上的评估表明,Auto-RecSys显著减少了每个实验周期所需的人工时间,并随着剧本的成熟提高了执行可靠性。
英文摘要
Auto-research agents have shown the potential to automate hypothesis generation, experiment execution, and iterative refinement. However, scaling this paradigm to industry-scale recommendation models introduces two challenges: (1) long feedback loops, where model training can take days, making serial iteration prohibitively slow and requiring parallel exploration across multiple research directions; and (2) system complexity, where large configurations, fragile infrastructure dependencies, and multi-day GPU jobs require robust and recoverable execution. We present Auto-RecSys, an autonomous research system for long-horizon experimentation on industry-scale recommendation models. Auto-RecSys addresses these challenges through three harness designs: (1) distributed asynchronous execution for running multiple experiments in parallel across servers, (2) centralized cross-server memory for persistent and recoverable execution across sessions and failures, and (3) cognitive-procedural separation, where natural-language skill files guide LLM reasoning while deterministic scripts enforce operational correctness. Auto-RecSys further employs a dual-loop self-evolving architecture: an Execution Evolution Loop in which model-specific playbooks accumulate operational knowledge by recording failed attempts and crystallizing successful pipelines, and an Idea Evolution Loop in which experimental outcomes inform subsequent ideation. Evaluated on recommendation models, Auto-RecSys significantly reduces the human time required per experiment cycle and improves execution reliability as its playbooks mature.
Comments16 pages, 4 figures