发表机构
N. D. Zelinsky Institute of Organic Chemistry RAS; National Research University Higher School of Economics; Department of Chemistry and Biochemistry, Florida State University(俄罗斯科学院泽林斯基有机化学研究所; 国立研究高等经济学院; 佛罗里达州立大学化学与生物化学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究测试了前沿大语言模型对分子三维结构的理解能力,发现新模型在构象排序上接近力场水平,且该能力与科学推理基准表现相关,为AI分子设计提供了新起点。
AI 中文摘要
大语言模型(LLMs)已展现出解决以自然语言表达的复杂化学问题的强大能力。然而,化学家执行的许多任务需要理解化合物的三维结构并对其进行推理——据我们所知,这些能力既非LLM开发者的初衷,也未在当代模型中得到测试。这种理解对于自主分子发现和药物开发至关重要,而这正是人工智能应用于化学领域的圣杯之一。我们在27个小有机分子的810个几何构型和DFT能量上测试了现代LLM的这一能力。引人注目的是,虽然2025年7月之前发布的模型在按稳定性对构象异构体排序方面表现挣扎,但许多近期前沿模型,包括GPT-5.6 Sol、Kimi K3和Gemini 3.6 Flash,取得了具有竞争力的准确性,超越了通用力场(Universal Force Field),而GPT-6 Astra和Claude Opus 5则非常接近现代的GFN-FF力场。对模型排序解释的分析表明,它们对具有明确分子内相互作用(如氢键)的分子表现最佳,这些相互作用在科学语言中得到了很好的量化;同时,除两个最新模型(GPT-6 Astra和Claude Opus 5)外,所有模型在处理由松散定义概念(如环张力)支配的分子时表现挣扎,而力场在这些方面表现出色,这表明科学语言本身可能对LLM施加了限制。值得注意的是,模型对构象异构体排序的能力与其在科学、编码和抽象推理基准上的表现密切相关,这表明该能力是从模型在相互关联但概念不同的任务上的训练中意外涌现的。得益于这种将单纯科学知识泛化为分子结构理解的能力,当前前沿模型为AI驱动的分子设计提供了一个现实的起点。
英文摘要
Large language models (LLMs) have already shown strong capabilities in solving complex chemical problems expressed in natural language. Yet, many tasks performed by chemists require understanding of 3D structures of chemical compounds and reasoning about them -- abilities, which, to the best of our knowledge, were neither intended by LLM developers nor tested in contemporary models. Such understanding is essential for autonomous molecular discovery and drug development, which is one of the Holy Grails of AI application to chemistry. We tested this capability in modern LLMs on 810 geometries and DFT energies of 27 small organic molecules. Strikingly, while models released before July 2025 struggled to rank conformers by stability, many recent frontier models, including GPT-5.6 Sol, Kimi K3, and Gemini 3.6 Flash achieved competitive accuracy, outperforming the Universal Force Field, and GPT-6 Astra and Claude Opus 5 came very close to a modern GFN-FF force field. Analysis of models' explanations for their rankings suggests that they perform best for molecules with well-defined intramolecular interactions -- hydrogen bonds -- which are well quantified in scientific language; at the same time all models except the two newest -- GPT-6 Astra and Claude Opus 5 -- struggle with molecules governed by loosely defined concepts (e.g., ring strain), where force fields excel, suggesting that the scientific language itself might impose constraints on LLMs. Notably, a model's ability to rank conformers is strongly associated with its performance on scientific, coding, and abstract-reasoning benchmarks, suggesting that it emerged unintendedly from models training on linked but conceptually different tasks. Thanks to this generalization of mere scientific knowledge into understanding molecular structures, current frontier models offer a realistic starting point for AI-driven molecular design.
Comments28 pages, 3 figures, 6 extended data figures. Supplementary Information included. Code and data available at https://github.com/TheorChemGroup/LLMConfBench