发表机构
MIT(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究自主研究中解决方案搜索过程的效率问题,提出用帕累托前沿曲线下面积评估AR系统,比较多种搜索算法,发现无单一最优结构,且效率与结果质量是不同维度,引入流体搜索,该方法在评估任务中总体搜索效率最高。
AI 中文摘要
人工智能驱动的自主研究(AR)系统在广泛任务中日益有效,但其性能仍主要按最终结果质量评估。本文认为解决方案搜索过程的效率是同样重要却常被忽视的性能维度。随着AR从数学和编码等低成本验证领域扩展到需昂贵物理实验评估解决方案的现实科学环境,搜索效率愈发重要。为考量此维度,我们提议用帕累托前沿曲线下面积(AUC)评估AR系统,同时兼顾最终结果质量。我们比较了包括爬山法、束搜索、树搜索和进化搜索在内的几类搜索算法在十二个系统优化任务中的表现。发现没有单一搜索结构始终最有效,且搜索效率和最终结果质量是不同性能维度。因最有效的搜索策略通常事先未知,我们引入名为流体搜索的自适应程序,它使用组合式策略在搜索过程森林中动态分配固定评估预算。在评估任务中,流体搜索实现了最高总体搜索效率,与提前为每个任务给定最佳搜索结构的任务特定预言机性能相近。
英文摘要
AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, is still evaluated primarily by the quality of the final outcome. In this paper, we argue that the efficiency of the solution-search process is an equally important but often overlooked dimension of performance. A strong AR system should not only produce high-quality results, but also reach them with as small a budget as possible. Search efficiency will become increasingly important as AR expands from domains with inexpensive verification, such as mathematics and coding, to real-world scientific settings in which solution evaluation may require costly physical experiments. To capture this dimension, we propose evaluating AR systems using the area under the curve (AUC) of the Pareto frontier, alongside final outcome quality. We compare several families of search algorithms, including hill climbing, beam search, tree search, and evolutionary search, across twelve systems-optimization tasks. We find that no single search structure is consistently the most efficient. We also show that search efficiency and final outcome quality are distinct performance dimensions: a method that eventually achieves the best result may nevertheless improve slowly and consume substantially more evaluation budget before reaching that result. Because the most effective search policy is generally unknown in advance, we introduce an adaptive procedure called fluid search, which uses a portfolio bandit to dynamically allocate a fixed evaluation budget across a forest of search processes. Across the evaluated tasks, fluid search achieves the highest overall search efficiency, closely matching the performance of a per-task oracle that is given the best search structure for each task in advance.