发表机构
Yusuf Hamied Department of Chemistry, University of Cambridge; John A. Paulson School of Engineering and Applied Sciences, Harvard University; Department of Materials Science and Engineering, Massachusetts Institute of Technology; Department of Molecular Science and Software Engineering, University of California, Berkeley; Sandia National Laboratories(剑桥大学; 哈佛大学; 麻省理工学院; 加州大学伯克利分校; 桑迪亚国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出基于NequIP和Allegro架构的等变MLIP基础势,解决速度与准确性的权衡问题,兼具高推理速度、强可扩展性及出色准确性,可降低超大规模数据集训练的计算成本。
AI 中文摘要
机器学习原子间势(MLIPs)已成为计算材料科学与化学领域的变革性工具,基于大规模多样数据集训练的通用势(即基础模型),现常被用于目标化学空间的下游微调。由此产生的模型(如分子动力学(MD))的诸多科学应用,要求具备高推理速度、训练速度及准确性。本研究探究等变MLIPs的极限——这类模型架构直接编码物理对称性,以实现上述相互竞争的目标,尤其在数据效率不那么关键的超大规模数据集场景下。我们展示了如何解决该权衡问题,并提出基于NequIP和Allegro等变MLIP架构的系列基础势,其在社区基准测试(涵盖材料发现、热导率预测、近平衡力学与热力学性质)中,兼具领先的推理速度、强可扩展性及出色准确性。NequIP架构内实现的加速技术,如今可在超大规模数据集上训练高精度基础势,计算成本大幅降低。此外,我们指出提升材料发现模型准确性的工作应聚焦于数据集多样性,以及对过渡金属化合物能量表面的改进、一致描述。
英文摘要
Machine-learned interatomic potentials (MLIPs) have emerged as a transformative tool for computational materials science and chemistry, with universal potentials trained on large and diverse datasets now routinely deployed as 'foundation models' for downstream fine-tuning in targeted chemical spaces. Many scientific applications of the resulting models, such as molecular dynamics (MD), require high inference and training speeds as well as accuracy. In this work, we examine the limits of equivariant MLIPs, which directly encode physical symmetries in model architectures, to achieve these competing targets -- particularly in the regime of extremely large datasets where data efficiency is less critical. We show how this trade-off can be addressed, and present a family of foundation potentials in the NequIP and Allegro equivariant MLIP architectures which achieve leading inference speeds and strong scalability as well as excellent accuracies across a range of community benchmarks -- spanning materials discovery, thermal conductivity prediction, and near-equilibrium mechanical and thermodynamic properties. Accelerations implemented within the NequIP infrastructure now permit training of high-accuracy foundation potentials on ultra-large datasets with dramatically reduced computational cost. Alongside, we show that efforts to improve model accuracy for materials discovery should focus on dataset diversity and improved, consistent descriptions of transition metal compound energy surfaces.