arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RxnOptBench:用于有机方法学中反应条件优化的LLM基准测试

RxnOptBench: Benchmarking LLMs for Reaction-Condition Optimization in Organic Methodology

Lingli Ge, Yubin Wang, Junyuan Gao, Jiahe Song, Jiaxing Sun, Boyu Zhu, Haote Yang, Jingchao Wang, Lixin Ma, Jiang Wu, Yuqiang Li, Conghui He

arXiv 2610.02242首次发表:更新:

发表机构

Shanghai Jiao Tong University; Shanghai Innovation Institute; Shanghai Artificial Intelligence Laboratory(上海交通大学; 上海创新研究院; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RxnOptBench是一个新基准,用于评估LLM在有机方法学中基于真实湿实验数据优化反应条件的能力,发现即使最佳模型仍有显著提升空间。

AI 中文摘要

化学反应条件优化——选择能够共同最大化产率和立体选择性的催化剂、配体、溶剂、试剂、温度、时间和气氛——是有机方法学研究中的一个核心且依赖判断的子任务,大型语言模型(LLMs)被日益期望能够支持这一任务。然而,现有的化学基准测试评估的是反应类别标注、逆合成或SMILES操作,并不要求模型阅读真实的条件筛选表格并选出最佳方案。我们提出了RxnOptBench,这是一个基准测试,其每个选项和先例均来自2025年发表的有机方法学论文优化表格中的真实湿实验条目,通过一个由声明的主要效用导出的连续相对分数进行评分,该效用结合了报告的产率与对映体过量(ee)、非对映体比率(dr)和区域异构体比率(rr),并配备了一个配对的“有先例vs无先例”设计,以将文献证据的上下文内使用与参数记忆区分开来。在九个前沿LLM和三个化学LLM的评估中,即使是最好的模型也留下了很大的提升空间:化学专业模型在多轴选择上降至随机基线水平,而开放权重模型已缩小了与专有前沿模型的大部分差距。我们发布了最终经人工审核的基准测试集和评估代码。

英文摘要

Chemical reaction-condition optimization -- choosing the catalyst, ligand, solvent, reagent, temperature, time, and atmosphere that jointly maximize yield and stereoselectivity -- is a central, judgement-laden subtask of organic methodology research that large language models are increasingly expected to support. Yet existing chemistry benchmarks evaluate reaction-class labelling, retrosynthesis, or SMILES manipulation, and do not ask models to read a real condition-screening table and pick the best set. We introduce RxnOptBench, a benchmark whose every option and precedent is a real wet-lab entry mined from the optimization tables of organic-methodology papers published in 2025, graded by a continuous relative score derived from a declared headline utility that combines reported yield with enantiomeric excess (ee), diastereomeric ratio (dr), and regioisomeric ratio (rr), and equipped with a paired precedents-vs-no-precedents design that isolates in-context use of literature evidence from parametric memorization. Across nine frontier LLMs and three Chemistry LLMs, even the best models leave substantial headroom: chemistry-specialized models fall to the random-baseline floor on multi-axis selection, while open-weight models have closed most of the gap to proprietary frontier models. We release the final human-reviewed benchmark test set and evaluation code.

CommentsAccepted to NeurIPS 2026 (Evaluations & Datasets Track)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑