arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TSBench:一个基于物理的基准,用于评估大语言模型对化学反应机理的理解

TSBench: A physics-grounded benchmark for evaluating LLM understanding of chemical reaction mechanisms

Xiaohu Xu, Tong Zhu

arXiv 2609.08503首次发表:更新:

AI 中文总结

TSBench通过量子化学验证的过渡态构建任务,评估七个大语言模型在78个基元反应上的机理理解,成功率从50.4%提升至66.8%,揭示了模型在复杂反应中的局限。

AI 中文摘要

理解一个化学反应需要将符号化的反应物-产物描述映射到原子重排所经过的三维路径上,然而针对大语言模型(LLMs)的化学基准大多只考察事实性知识和基于文本的推理。在此,我们引入TSBench,这是一个基准测试,其中LLM智能体使用结构编辑工具构建三维过渡态(TS)猜测,并由自动化量子化学流水线进行验证,从而得出基于物理的通过/失败判定。在对78个基元反应上七个前沿LLM的546次评估中,在诊断驱动的修订下,总体成功率从50.4%提升至66.8%;最好的模型在最简单的反应上接近90%的成功率,但性能随机理复杂度的增加而急剧下降。大多数失败的尝试产生了局部看似合理的鞍点,其反应路径却导向了错误的反应物-产物对,这揭示了模型具备局部几何直觉,但对整体反应坐标缺乏稳健的把握。TSBench为LLM智能体在诸如合成规划和自主实验等对机理敏感的任务中建立了一个机理层面的衡量标准。

英文摘要

Understanding a chemical reaction requires mapping a symbolic reactant-product description onto the three-dimensional pathway through which atoms rearrange, yet chemistry benchmarks for large language models (LLMs) largely probe factual knowledge and text-based reasoning. Here we introduce TSBench, a benchmark in which an LLM agent uses structure-editing tools to construct three-dimensional transition-state (TS) guesses verified by an automated quantum-chemical pipeline, yielding a physics-grounded pass/fail verdict. Across 546 evaluations of seven frontier LLMs on 78 elementary reactions, the aggregate success rate rose from 50.4% to 66.8% under diagnosis-driven revision; the best models approached 90% on the simplest reactions, yet performance dropped sharply with mechanistic complexity. Most failed attempts produced locally plausible saddle points whose reaction paths led to the wrong reactant-product pair, revealing local geometric intuition without a robust grasp of the global reaction coordinate. TSBench establishes a mechanism-level yardstick for LLM agents in mechanism-sensitive tasks such as synthesis planning and autonomous experimentation.

Comments35 pages, including 13 pages of supplementary information

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑