AI 中文总结
该研究针对自主机器学习工程的长时任务,提出信息范式并实现为Iris智能体,其在MLE-Bench上12小时预算下获64.9%任意奖牌率,还具备跨领域泛化能力。
AI 中文摘要
机器学习工程等长时自主研究任务要求系统在有限预算下做出相互依赖的决策。现有基于大语言模型(LLM)的智能体通常通过树、图或链结构组织候选解决方案的改进,即搜索过程决定信息的获取与管理方式,我们将这种设计称为以解决方案为中心的搜索。本文提出信息范式,其中不断演化的信息状态代表系统对任务的理解,并指导解决方案的改进。我们将该范式实例化为Iris,一种查询-修正循环:在信息获取方面,Iris从当前信息状态生成本地行动计划,并使用认知行动探查决策关键未知项,且不修改已保留的解决方案;在信息管理方面,Iris将实验中的观察结果合成为由可修正声明组成的任务知识,这些声明具有明确的范围和状态,它会随着新证据的到来更新该知识,并根据所需的详细程度从原始证据、结构化摘要或任务知识中构建每个决策上下文。在MLE-Bench上,Iris在12小时预算下达到64.9%的任意奖牌率,是对比系统中最高的;在涵盖工具工程和模型后训练的四项任务中,Iris还展现出跨领域泛化能力。
英文摘要
Long-horizon autonomous research tasks such as machine learning engineering require systems to make interdependent decisions under a limited budget. Existing LLM-based agents typically organize candidate-solution improvement through tree, graph, or chain structures, meaning that the search process determines how information is acquired and managed. We call this design solution-centric search and propose instead the information paradigm, in which an evolving information state represents the system's understanding of the task and guides solution improvement. We instantiate this paradigm in Iris, an inquiry-revision loop. For information acquisition, Iris generates local action plans from the current information state and uses epistemic actions to probe decision-critical unknowns without modifying the retained solution. For information management, Iris synthesizes observations across experiments into task knowledge composed of revisable claims with explicit scope and status. It updates this knowledge as new evidence arrives and constructs each decision context from raw evidence, structured summaries, or task knowledge at the required level of detail. On MLE-Bench, Iris attains a 64.9% any-medal rate under a 12-hour budget, the highest among compared systems. Across four tasks spanning harness engineering and model post-training, Iris also demonstrates cross-domain generalization.