发表机构
School of Computer Science, Carnegie Mellon University; DP Technology(计算机科学系,卡内基梅隆大学; DP科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究材料预测中人工智能代理的工作流程,通过多方面搜索产生评估更改,经留出数据测试,发现不同建模机制下的有效变化,证明闭环代理决策及代码可复用组合,还提供了评估设计。
AI 中文摘要
人工智能研究代理可能在未找到适用于新材料的建模变化时提高其所见分数。我们提出了一个更严格的问题:经过反复实验后,所选变化在从未进入循环的数据上是否依然有效,其代码能否复用?我们将搜索分为对特征、模型、表示和训练数据的更改。七个搜索在十个Matbench端点上产生了701个评估更改。代理仅接收五个内部折叠的平均值,减少对任何单个开发分割的依赖。然后冻结所选代码并在未触及的留出数据上进行一次评估。十个选择中有九个仍然是测试过的最佳单一干预措施。幸存的变化揭示了两种材料建模机制。仅对于成分,特征、模型和表示的变化提供了可比的改进途径。对于结构任务,更丰富的几何描述符以及模型或校准变化可降低平均留出MAE。这些结果表明,在材料预测中,闭环代理可以做出经得起未见证据检验的决策,并且代码更改可跨任务复用和组合。更广泛地说,它们为测试反馈循环之外的可执行发现提供了一种评估设计。
英文摘要
Auto Research uses language-model agents to propose, implement, and evaluate machine-learning changes in a closed loop, but is usually judged by its terminal pipeline. A terminal score cannot reveal which technical decision produced a gain or distinguish a reusable discovery from a change adapted to development feedback. We introduce intervention-centered Auto Research, which validates research decisions rather than only final artifacts and makes their reliability measurable. Feature, Model, Representation, and Data axes are searched independently with inner five-fold feedback. Each axis winner is frozen before an outer-holdout matrix compares all alternatives on evidence the loop never sees. Across 701 agent-executed attempts spanning ten Matbench endpoints, outer evidence confirms the selected intervention on nine of ten endpoints and preserves 89.3\% of non-tied intervention orderings. It also rejects an aggregate Representation gain that inner feedback endorsed. The resulting matrix reveals an information-dependent hierarchy. Composition-only tasks support several routes to improvement, whereas structure-informed tasks favor local geometry features and complementary tree ensembles. A subsequent compatibility test combines already frozen Feature and Model code without further search or tuning and raises mean outer-holdout improvement from 19.0\% to 26.3\%. By validating decisions rather than only artifacts, this design turns adaptive search into reusable evidence wherever agents propose executable alternatives against a fixed evaluator.