发表机构
Princeton University; University of Pittsburgh; Northeastern University; University of Florida(普林斯顿大学; 匹兹堡大学; 东北大学; 佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究利用大语言模型学习化学反应机理推理,构建大规模推理数据集及福山基准,通过机理感知训练微调Qwen3 - 30B - A3B,在福山基准集A上精确路径匹配超FlowER模型,增强了语言模型的化学推理能力。
AI 中文摘要
反应机理由解释化学转化的基本反应的逐步序列组成。因此,学习机理逻辑对于增强大语言模型(LLMs)的基本化学智能至关重要。反应机理的逐步推导与推理LLMs的推理范式自然契合。然而,当前化学LLMs主要强调用于产物预测和逆合成的粗粒度名称反应,常导致物理不一致和幻觉。相比之下,用于机理推断的专门小规模生成模型通常在不同化学空间中泛化能力受限。为克服这些限制,我们构建了一个新颖的大规模反应机理推理数据集。此外,我们建立了福山基准,这是一个源自福山《高等有机反应机理》一书的具有挑战性的基准,用于严格评估模型在分层机理推理上的性能。我们微调后的Qwen3 - 30B - A3B在福山基准集A上实现了8.3%的精确路径匹配,超过了专门的FlowER模型(5.1%),表明机理感知训练显著增强了语言模型中的化学推理能力。
英文摘要
Reaction mechanisms consist of the step-by-step sequences of elementary reactions that explain chemical transformations. Learning the mechanism logic is therefore essential for enhancing the fundamental chemical intelligence of large language models (LLMs). The stepwise deduction of reaction mechanism aligns naturally with the reasoning paradigms of reasoning LLMs. However, current chemical LLMs primarily emphasize coarse-grained name reactions for product prediction and retrosynthesis, often leading to physical inconsistencies and hallucinations. In contrast, specialized small-scale generative models for mechanism inference typically suffer from restricted generalization capacity across diverse chemical spaces. To overcome these limitations, we built a novel, large-scale reasoning dataset of reaction mechanisms. Furthermore, we established the FukuyamaBench, a difficult benchmark derived from Fukuyama's Advanced Organic Reaction Mechanism book, to rigorously evaluate model performance on hierarchical mechanism reasoning. Our fine-tuned Qwen3-30B-A3B achieves 8.3% exact pathway match on FukuyamaBench Set~A, surpassing the specialized FlowER model (5.1%), demonstrating that mechanism-aware training substantially enhances chemical reasoning in language models.