arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自动化发现没有普遍优越的框架

Automated Discovery Has No Universally Superior Harness

Akshat Gupta, Jermaine Lei, Alexander Lu, Gopala Anumanchipalli, Leshem Choshen

arXiv 2607.18235首次发表:更新:

AI 中文总结

研究自主发现系统框架,通过系统分解和评估发现其存在泛化问题,无固定框架普遍优越,OpenEvolve变体表现不佳。利用早期发现进展预测性能,提出自适应分配实验,优于固定和非自适应选择,推动从固定框架选择转向在线适应。

AI 中文摘要

诸如OpenEvolve和TTT-Discover等自主发现系统常被用作通用框架。但实际上它们是复合系统,将多种关于存档、亲本选择、探索和预算分配的设计选择组合成单一方案。由于发现运行成本高且具有内在随机性,现有框架常因独立试验太少难以区分关键方法改进和运行差异。我们系统分解相关搜索框架并评估。结果表明发现框架存在泛化问题,没有固定框架在所有评估模型问题对中都可靠优越,OpenEvolve变体通常表现不如更简单替代方案。早期发现进展可预测最终性能,据此提出预算匹配的自适应分配实验,优于固定框架选择和非自适应框架集合。这些结果促使从固定框架选择转向由早期性能引导的在线适应。我们发布所有运行池及基线空分布作为可重复使用的统计基础设施。

英文摘要

Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However, in practice these are composite systems combining several design choices about archives, parent selection, exploration, and budget allocation into a single recipe. Because discovery runs are expensive and inherently stochastic, existing harnesses are often compared using too few independent trials to distinguish key methodological improvements from run-to-run variance. We systematically decompose OpenEvolve-style evolutionary search and the TTT-Discover search harness into its constituent components and systematically evaluate 30 budget-matched harnesses across 12 model-problem pairs using more than 3.1 million LLM rollouts and repeated-trial statistical analysis. Our results show that discovery harnesses have a generalization problem: No fixed harness is reliably superior across the evaluated model-problem pairs, and variants of OpenEvolve generally underperform simpler alternatives. Thus, harness choice is better viewed as a hyperparameter rather than as a universal recipe, and should be tailored to the specific problem and underlying model. We also find that early discovery progress predicts final performance, and use this property to present a budget-matched adaptive-allocation experiment that starts multiple harnesses, prunes weak partial runs, and reallocates compute to stronger survivors, outperforming both commitment to a randomly sampled fixed harness and a non-adaptive harness ensemble. Together, these results motivate shifting from fixed harness selection to online adaptation guided by early performance. We release all run pools including baseline null distributions for every model-problem pair as reusable statistical infrastructure against for future harness proposals.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑