发表机构
Shanghai Jiao Tong University; Shanghai Innovation Institute; Shanghai Artificial Intelligence Laboratory(上海交通大学; 上海创新研究院; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RxnOptBench是一个新基准,用于评估LLM在有机方法学中基于真实湿实验数据优化反应条件的能力,发现即使最佳模型仍有显著提升空间。
AI 中文摘要
化学反应条件优化——选择能够共同最大化产率和立体选择性的催化剂、配体、溶剂、试剂、温度、时间和气氛——是有机方法学研究中的一个核心且依赖判断的子任务,大型语言模型(LLMs)被日益期望能够支持这一任务。然而,现有的化学基准测试评估的是反应类别标注、逆合成或SMILES操作,并不要求模型阅读真实的条件筛选表格并选出最佳方案。我们提出了RxnOptBench,这是一个基准测试,其每个选项和先例均来自2025年发表的有机方法学论文优化表格中的真实湿实验条目,通过一个由声明的主要效用导出的连续相对分数进行评分,该效用结合了报告的产率与对映体过量(ee)、非对映体比率(dr)和区域异构体比率(rr),并配备了一个配对的“有先例vs无先例”设计,以将文献证据的上下文内使用与参数记忆区分开来。在九个前沿LLM和三个化学LLM的评估中,即使是最好的模型也留下了很大的提升空间:化学专业模型在多轴选择上降至随机基线水平,而开放权重模型已缩小了与专有前沿模型的大部分差距。我们发布了最终经人工审核的基准测试集和评估代码。
英文摘要
Chemical reaction-condition optimization -- choosing the catalyst, ligand, solvent, reagent, temperature, time, and atmosphere that jointly maximize yield and stereoselectivity -- is a central, judgement-laden subtask of organic methodology research that large language models are increasingly expected to support. Yet existing chemistry benchmarks evaluate reaction-class labelling, retrosynthesis, or SMILES manipulation, and do not ask models to read a real condition-screening table and pick the best set. We introduce RxnOptBench, a benchmark whose every option and precedent is a real wet-lab entry mined from the optimization tables of organic-methodology papers published in 2025, graded by a continuous relative score derived from a declared headline utility that combines reported yield with enantiomeric excess (ee), diastereomeric ratio (dr), and regioisomeric ratio (rr), and equipped with a paired precedents-vs-no-precedents design that isolates in-context use of literature evidence from parametric memorization. Across nine frontier LLMs and three Chemistry LLMs, even the best models leave substantial headroom: chemistry-specialized models fall to the random-baseline floor on multi-axis selection, while open-weight models have closed most of the gap to proprietary frontier models. We release the final human-reviewed benchmark test set and evaluation code.
CommentsAccepted to NeurIPS 2026 (Evaluations & Datasets Track)