arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在线代理修复:在闭环发现中将高保真反馈与搜索长度解耦

Online Surrogate Repair: Decoupling High-Fidelity Feedback from Search Length in Closed-Loop Discovery

Xiaotang Feng, Philip Torr, Bruno Andreis

arXiv 2609.07655首次发表:更新:

发表机构

University of Oxford; Slater Labs(牛津大学; Slater实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出在线代理修复(OSR),在闭环发现中利用稀疏高保真评估更新代理,将反馈频率与搜索长度解耦,显著减少oracle查询,提升优化效率。

AI 中文摘要

闭环人工智能科学家能够以较低的边际计算成本生成候选设计,而可靠的反馈可能需要湿实验室合成、表征或高保真计算。通过定制实验室自动化来解决这种不平衡仍然需要大量基础设施且成本高昂,而用固定代理替换新实验会留下持续的模型误差,这些误差可能被优化放大。我们提出了在线代理修复(OSR),一种闭环算法,它使用稀疏的高保真评估来更新代理,而代理搜索过程主要使用廉价的代理反馈进行。采集规则选择代理累积提案中的哪些设计接受高保真评估,所得标签更新后续情节中使用的代理。在受控的合成环境中,我们证明改善全局代理拟合并不一定能减少最大遗憾,而Q90-UCB和期望改进(EI)通过将评估引导至决定优化器决策的区域,显著减少了遗憾。在MADE上,每个情节后接受高保真反馈的对照组需要6.36–7.23倍更多的oracle查询才能匹配两种LLM编排器下的在线EI,而在非LLM的Chemeleon+MLIP工作流下则需要10.27倍更多。在线代理修复引入了介于固定代理操作和每个情节后高保真反馈之间的新颖的第三种反馈机制,将高保真评估的频率与代理搜索的持续时间分离开来。

英文摘要

Closed-loop AI scientists can generate candidate designs at low marginal computational cost, whereas reliable feedback may require wet-lab synthesis, characterization, or high-fidelity computation. Addressing this imbalance through custom laboratory automation remains infrastructure-intensive and costly, while replacing new experiments with a fixed surrogate leaves persistent model errors that can be amplified by optimization. We propose \emph{online surrogate repair} (OSR), a closed-loop algorithm that uses sparse high-fidelity evaluations to update the surrogate throughout a longer agent search conducted primarily with inexpensive surrogate feedback. An acquisition rule selects which designs from the agent's accumulated proposals receive high-fidelity evaluation, and the resulting labels update the surrogate used in subsequent episodes. Across controlled synthetic environments, we demonstrate that improving global surrogate fit does not necessarily reduce maximum regret, whereas Q90-UCB and expected improvement (EI) substantially reduce regret by directing evaluations toward regions that determine the optimizer's decisions. On MADE, controls receiving high-fidelity feedback after every episode require $6.36$--$7.23\times$ more oracle queries to match Online EI under two LLM orchestrators and $10.27\times$ more under the non-LLM Chemeleon+MLIP workflow. Online surrogate repair introduces a novel third feedback regime between fixed-surrogate operation and high-fidelity feedback after every episode, separating the frequency of high-fidelity evaluation from the duration of the agent's search.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑