发表机构
University of California San Diego(加利福尼亚大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出AIBuildAI-2.5智能体系统,通过LLM引导的树搜索、资源感知调度和成本感知路由,解决自主AI模型开发中的效率问题,在MLE-Bench上以73.3%奖牌率排名第一。
AI 中文摘要
能够自动构建人工智能(AI)模型的自主智能体可拓宽科学和工程领域对AI的访问。此类智能体的一种流行范式将模型构建视为代码搜索问题,并通过树搜索来解决,其中每个节点是一个候选程序,树通过从父程序生成子程序来生长,这些智能体现在在现实基准上已接近经验丰富的AI工程师的能力。然而,这些智能体在效率方面存在三个尚未完全解决的弱点。首先,在现实预算内只能执行少量候选程序,因此按执行奖励对节点排序的搜索规则(如蒙特卡洛式树搜索)依赖少量且噪声大的分数,导致选择下一个要探索的节点时效果较差。其次,没有采用资源感知策略来调度训练作业,这可能会降低硬件利用率和训练效率。第三,每次智能体调用都由一个强大的模型提供服务,这增加了推理成本。在此,我们引入AIBuildAI-2.5,一个使用LLM智能体执行树搜索并解决上述三个问题的智能体系统。AIBuildAI-2.5提出了一种新颖的LLM引导的树搜索,其中评判器根据每个候选的预期改进、依据性和可行性进行评分,选择器则根据这些分数和搜索状态对候选池进行排序。此外,AIBuildAI-2.5包含一个调度器,在考虑当前硬件资源状态的情况下启动训练作业,以及一个路由器,将成本较低的LLM分配给要求较低的任务,同时保留最强大的LLM用于AI模型构建工作流中最具挑战性的子任务。AIBuildAI-2.5在MLE-Bench上以73.3%的奖牌率排名第一,并在来自AIRS-Bench的六个自主AI研究任务上优于强基线。
英文摘要
Autonomous agents that automatically build artificial intelligence (AI) models could broaden access to AI across science and engineering. A popular line of such agents frames model building as a code search problem and solves it by tree search, in which each node is a candidate program and the tree grows by generating a child program from a parent, and these agents now approach the capability of experienced AI engineers on realistic benchmarks. However, these agents have three weaknesses in efficiency that have not been fully addressed. First, only a small number of candidates can be executed within a realistic budget, so search rules that rank nodes by executed rewards, such as Monte Carlo-style tree search, rely on few and noisy scores and select the next node to explore less effectively. Second, no resource-aware strategy is used to schedule training jobs, which can lower hardware utilization and training efficiency. Third, every agent call is served by a single powerful model, which inflates inference cost. Here we introduce AIBuildAI-2.5, an agentic system that carries out the tree search with LLM agents and addresses each of the three issues. AIBuildAI-2.5 proposes a novel LLM-guided tree search, in which a judge scores each candidate on its expected improvement, grounding, and feasibility, and a selector ranks the pool of candidates from these scores and the state of the search. In addition, AIBuildAI-2.5 comprises a scheduler that launches training jobs with the current hardware resource status taken into account and a router that assigns lower-cost LLMs to less demanding tasks while reserving the most capable LLM for the most challenging sub-tasks in the AI model building workflow. AIBuildAI-2.5 ranks first on MLE-Bench with a medal rate of 73.3%, and outperforms a strong baseline on six autonomous AI research tasks from AIRS-Bench.