AI 中文总结
该研究量化了将通用机器学习原子间势转换为材料特异性势所需的从头算分子动力学数据量,对比了不同框架的表现,提出了实用构建指南,还发现模型平均可提升稀缺数据下的预测性能。
AI 中文摘要
预训练的机器学习原子间势(即通用模型或基础模型)为原子级模拟提供了极具吸引力的起点,但在缺乏额外参考数据(微调)的情况下,其对材料特异性可观测量的精度往往有限。本文系统量化了将通用模型转换为从头算精度的材料特异性势所需的第一性原理数据量,并探讨微调是否一定优于从头训练。我们在包含稀有和反应事件的7种化学多样体系中,对比了5种通用MLIP框架:MACE-MP-0、SevenNet-0、GRACE-1L-OAM、MatterSim-v1-5M和ORB-v2。仅对10个AIMD导出的构型进行微调,对所研究体系而言不足;200个构型在有利情况下可成功,但结果仍高度依赖体系。相比之下,2000个AIMD构型构成了稳健的默认值,可产生低力和能量误差,并重现目标材料特异性可观测量。对AIMD轨迹进行适度密集的子采样可将所需轨迹长度缩短10倍,且模型质量损失很小。在相同数据集上从头训练,对MACE和SevenNet而言与简单微调具有竞争力,且通常精度略高,而GRACE需要更多数据。MoS₂中硫空位跳跃的能量曲线表明,低轨迹级误差不能保证正确的反应曲线,凸显了可观测量级验证的必要性。最后,我们证明在稀缺数据 regime 中,对独立训练的模型取平均可提升预测性能,且无需额外第一性原理成本。这些结果共同为将有限的AIMD参考数据转换为近DFT精度、适用于纳秒级模拟的可靠材料特异性MLIP提供了实用指南。
英文摘要
Pretrained machine-learning interatomic potentials, so-called universal or foundation models offer an appealing starting point for atomistic simulations, but their accuracy for material-specific observables often remains limited without additional reference data (fine-tuning). Here, we systematically quantify how much first-principles data are required to convert universal models into ab initio-accurate material-specific potentials, and ask whether fine-tuning is necessarily preferable to training from scratch. We compare five universal MLIP frameworks, MACE-MP-0, SevenNet-0, GRACE-1L-OAM, MatterSim-v1-5M and ORB-v2, across seven chemically diverse systems incorporating rare and reactive events. Fine-tuning on only 10 AIMD-derived configurations is insufficient for the investigated systems; 200 configurations succeed in favorable cases, but the outcome remains strongly system-dependent. By contrast, 2000 AIMD configurations constitute a robust default, yielding low force and energy errors and reproducing the target material-specific observables. Moderately dense sub-sampling of the AIMD trajectory reduces the required trajectory length tenfold with little loss in model quality. Training from scratch on the same datasets is competitive with, and often slightly more accurate than, naive fine-tuning for MACE and SevenNet, whereas GRACE requires more data. The energy profile for a sulfur-vacancy jump in MoS$_2$ reveals that low trajectory-level errors do not guarantee a correct reaction profile, highlighting the need for observable-level validation. Finally, we show that averaging independently trained models improves predictions in scarce-data regimes at no additional first-principles cost. Together, these results provide practical guidelines for converting limited AIMD reference data into reliable material-specific MLIPs for nanosecond-timescale simulations at near-DFT accuracy.