arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20891physics.chem-ph

RECOB:化学与材料科学实验优化的可靠基准测试

RECOB: Reliable Benchmarking of Experimental Optimization in Chemistry and Materials Science

Zikai Xie, Jiaming Wan, Linjiang Chen

首次发表
浏览论文内容

中文总结 AI 辅助

RECOB是一个基于化学与材料科学物理实验数据构建的黑箱优化基准,包含16个任务,通过统一协议评估多种优化方法,HEBO和qNEHVI分别领先单目标和多目标,且排名在重训练与回放下保持稳定,为实验优化提供可靠测试平台。

中文摘要 AI 辅助

实验科学的优化方法通常在合成函数上进行评估,这些函数虽可复现,但忽略了真实实验的重要特征。我们引入了RECOB(REliable Chem Optimization Benchmark,可靠化学优化基准,GitHub仓库:此https URL),这是一个完全基于化学和材料科学物理实验生成数据构建的黑箱优化基准。该套件包含14个单目标和两个多目标任务,涵盖化学反应、材料配方、电化学系统、连续流过程和自动化实验室。每个任务都提供了其决策变量、可行域、物理约束、目标方向和实验来源的机器可读规范。可连续查询的学习型代理模型通过重复留出验证和预设准入标准进行筛选,而测量表回放则能利用原始实验响应进行评估。在统一的配对评估协议下,我们比较了十种单目标和八种多目标优化方法。HEBO在单目标综合排名中表现最佳,而qNEHVI在多目标比较中领先。基于模型的方法通常优于非自适应基线,尽管其计算开销差异显著。我们进一步通过独立的代理模型重训练和测量表回放来评估基准的可靠性。优化器的综合排名在重训练的代理模型间保持高度一致,而回放仅使用物理测量响应就保留了广泛的性能层级。这些结果共同表明,RECOB能够可复现地区分优化器性能,作为化学和材料科学中黑箱优化的一个基于实验且经过可靠性测试的基准。

英文摘要

Optimization methods for experimental science are often evaluated on synthetic functions that are reproducible but omit important characteristics of real experiments. We introduce RECOB (REliable Chem Optimization Benchmark, Github repository: \hyperlink{https://github.com/XieZikai/RECOB}{https://github.com/XieZikai/RECOB}), a black-box optimization benchmark constructed exclusively from data generated through physical experiments in chemistry and materials science. The suite contains 14 single-objective and two multi-objective tasks spanning chemical reactions, material formulations, electrochemical systems, continuous-flow processes, and automated laboratories. Each task provides a machine-readable specification of its decision variables, feasible domain, physical constraints, objective direction, and experimental provenance. Continuously queryable learned oracles are screened using repeated holdout validation and prespecified admission criteria, while measured-table replay enables evaluation using the original experimental responses. Under a common paired evaluation protocol, we compare ten single-objective and eight multi-objective optimization methods. HEBO achieves the best aggregate single-objective rank, while qNEHVI leads the multi-objective comparison. Model-based methods generally outperform non-adaptive baselines, although their computational overhead varies substantially. We further assess benchmark reliability using independent oracle retraining and measured-table replay. Aggregate optimizer rankings remain highly consistent across retrained oracles, while replay preserves the broad performance hierarchy using only physically measured responses. Together, these results show that RECOB can reproducibly distinguish optimizer performance as an experimentally grounded and reliability-tested benchmark for black-box optimization in chemistry and materials science.

发表机构

  • State Key Laboratory of Precision and Intelligent Chemistry, University of Science and Technology of China(中国科学技术大学精密与智能化学国家重点实验室)
  • Center for Scientific Intelligence Innovation(科学智能创新中心)
  • Department of Computer Science, Fudan University(复旦大学计算机科学与技术学院)
  • School of Chemistry, School of Computer Science, University of Birmingham(伯明翰大学化学学院、计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

↑