arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20753physics.chem-ph

基于机器学习分子电子密度的精确且可迁移的分子间势能

Accurate and Transferable Intermolecular Potential Based on Machine-Learned Molecular Electron Density

Dahvyd Wing, Mihail Bogojeski, Szabolcs Goger, Klaus-Robert Müller, Alexandre Tkatchenko

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出仅含4个通用参数的物理模型DensIP,在DES15K数据集上训练后,对含未训练分子的二聚体实现亚kcal/mol误差,长程相互作用性能优于先进通用MLFFs,可用于药物配体,为大规模生成精确合成训练数据提供新途径。

中文摘要 AI 辅助

机器学习力场(MLFFs)包含大量可学习参数,因此需要庞大的训练数据集。这为开发高精度、通用型MLFFs带来挑战,因为生成高质量的从头算参考数据计算成本高昂。经典经验势能提供了潜在廉价的合成训练数据来源,但现有模型常缺乏提供有用参考能量所需的精度。本文提出密度基分子间势能(DensIP),这是一种基于物理的分子间相互作用模型,使用机器学习得到的电子密度,仅含4个通用参数。我们在DES15K数据集上对DensIP进行训练和测试,该数据集包含小有机分子二聚体的CCSD(T)/CBS相互作用能。DensIP在包含训练集中未出现分子的二聚体(包括非平衡构象的分子)上实现了低于1 kcal/mol的误差,展现出强可迁移性。我们进一步证明DensIP可应用于药物配体这类大分子。值得注意的是,DensIP在长程相互作用方面优于最先进的通用型MLFFs,是大规模生成精确合成训练数据的有前景方法。

英文摘要

Machine-learned force fields (MLFFs) contain many learnable parameters and therefore require large training datasets. This poses a challenge for developing highly accurate, general-purpose MLFFs because generating high-quality ab initio reference data is computationally expensive. Classical empirical potentials offer a potentially inexpensive source of synthetic training data, but existing models often lack the accuracy needed to provide useful reference energies. Here, we introduce the density-based intermolecular potential (DensIP), a physics-based model of intermolecular interactions that uses machine-learned electron densities and only four universal parameters. We train and test DensIP on CCSD(T)/CBS interaction energies from DES15K, a dataset of dimers of small organic molecules. DensIP achieves sub-kcal/mol errors for dimers containing molecules absent from the training set, including molecules in non-equilibrium conformations, demonstrating strong transferability. We further show that DensIP can be applied to molecules as large as drug ligands. Notably, DensIP outperforms state-of-the-art general-purpose MLFFs for long-range interactions, making it a promising approach for generating accurate synthetic training data at scale.

补充信息

↑