arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种系统无关的元动力学反应训练数据生成策略:应用于气相有机反应

A System-Independent Metadynamics Strategy for Reactive Training Data: Application to Gas-Phase Organic Reactions

Wanrun Jiang, Jinzhe Zeng, Manyi Yang, Tong Zhu, Han Wang

arXiv 2609.29105首次发表:更新:

发表机构

AI for Science Institute; University of Science and Technology of China; Nanjing University; East China Normal University; Institute of Applied Physics and Computational Mathematics(人工智能科学研究所; 中国科学技术大学; 南京大学; 华东师范大学; 应用物理与计算数学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有气相有机反应MLIP训练数据局限于MEP附近的问题,提出基于局部RMSD集体变量的系统无关元动力学采样策略,生成含180万DFT标记构型的OpenRxn26数据集,使DPA3_rxn在偏离MEP轨迹上达到1.0 kcal/mol精度,验证了方法的通用性与效率。

AI 中文摘要

面向有机反应的通用机器学习原子间势(MLIPs)既需要在用于静态评估基本性质的最小能量路径(MEP)上保持准确性,也需要在用于模拟反应动力学的更广泛构型空间上保持准确性。现有的气相有机反应通用数据集依赖于准静态弛豫,将构型限制在MEP附近,因此在这些数据集上训练的模型可能在直接分子动力学轨迹上失效;这一差距是方法层面的,而非数据集规模问题。我们引入了一个时空分辨、系统无关的集体变量(CV):在随机划分的局部域内,相对于不断扩展的时间平均参考几何列表的笛卡尔RMSD。该CV驱动元动力学作为主要探索引擎,并辅以向过渡态(TS)的结构弛豫以增强TS周围的覆盖。在并发学习工作流中,这生成了OpenRxn26数据集,包含180万DFT标记的构型,覆盖H/C/N/O化学空间($N_{\mathrm{heavy}} \leq 30$)中的中性单重态单分子反应,包含社区数据集中代表性不足的反应性原子环境。基于OpenRxn26训练的DPA3模型(记为DPA3_rxn)在势垒高度和反应能上实现了可迁移的准确性。在偏离MEP的反应性轨迹上,DPA3_rxn是基准套件中唯一达到与标记方法相比1.0 kcal/mol能量准确性的模型,而静态基准上领先的领域MLIP(如MACE_OMol25)则下降数倍,表明准静态采样和MEP锚定基准不足以保证MLIP的动力学可靠性。因此,OpenRxn26为气相中性单重态有机反应提供了可直接用于分子动力学(MD)的反应训练数据,验证了该采样策略的通用性和效率。

英文摘要

General-purpose machine-learning interatomic potentials (MLIPs) for organic reactions need to be accurate on both the minimum energy path (MEP) for static evaluation of basic properties and the broader configurational space for simulating reaction dynamics. Existing general datasets for gas-phase organic reactions rely on quasi-static relaxation that confines configurations to the MEP vicinity, so models trained on them could fail on direct molecular-dynamics trajectories; the gap is methodological, not a question of dataset size. We introduce a spatiotemporally resolved, system-independent collective variable (CV): Cartesian RMSD within randomly partitioned local domains against an expanding list of time-averaged reference geometries. The CV drives metadynamics as the main exploration engine, supplemented by structural relaxation towards transition state (TS) to augment the coverage around TS. Within a concurrent-learning workflow, this produces OpenRxn26, a dataset of 1.8~M DFT-labeled configurations covering neutral singlet unimolecular reactions in the H/C/N/O chemical space ($N_\mathrm{heavy} \leq 30$), containing reactive atomic environments underrepresented in community datasets. Trained on OpenRxn26, a DPA3 model (denoted DPA3_rxn) achieves transferable accuracy on barrier heights and reaction energies. On off-MEP reactive trajectories, DPA3_rxn is the only model in the benchmark suite to reach 1.0 kcal/mol energy accuracy compared with the labeling method, where the domain MLIP leading on static benchmarks degrades several-fold (e.g. MACE_OMol25), showing the insufficiency of quasi-static sampling and MEP-anchored benchmarks for guaranteeing dynamics reliability of MLIPs. OpenRxn26 thus provides MD-ready reactive training data for gas-phase neutral singlet organic reactions, verifying the generality and efficiency of the sampling strategy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑