达到经验势速度的通用机器学习分子动力学方法
Universal Machine-learning Molecular Dynamics at the Speed of Empirical Potentials
浏览论文内容
中文总结 AI 辅助
研究人员提出兼顾精度与效率的等变势DPA4C,其多变体在吞吐量、误差及并行扩展性上表现优异,可实现接近第一性原理精度的通用分子动力学模拟,速度达经验势水平。
中文摘要 AI 辅助
目前尚无任何原子间势能同时兼顾跨化学体系的通用性、接近第一性原理的精度以及经验势的速度。本文提出DPA4C,这是一种等变势,其架构与压缩CUDA算子在部署约束下协同设计,以同时追求精度与效率。涵盖49倍参数范围的5种变体构成了所测精度-吞吐量前沿的高通量端;最大变体的精度接近MACE-Omat模型,测量吞吐量却高出约两个数量级;最紧凑变体在现有最快通用MLIP饱和吞吐量的1.92倍下,将能量、力和应力误差分别降低61.4%、48.1%和34.3%。所有5种变体均可在单GPU上完成百万原子级模拟,且在1024块16-GB NVIDIA V100 GPU上运行20.48亿原子的分子动力学时,弱缩放效率达83.3%-91.2%。与MEAM经验势相比,DPA4C-Nano在相同V100硬件上的单GPU扫描中,针对金刚石碳和FCC铜分别达到其1.8倍和2.5倍的饱和吞吐量。因此,DPA4C将量子训练的通用精度带入了此前仅与经验势相关的速度和体系规模区间。
英文摘要
No interatomic potential has offered universality across chemistry, near-first-principles accuracy and the speed of empirical potentials at once. Here we introduce DPA4C, an equivariant potential whose architecture and compressed CUDA operators are co-designed under deployment constraints to pursue accuracy and efficiency together. Five variants spanning a 49-fold parameter range form the high-throughput end of the measured accuracy--throughput frontier. The largest variant approaches the accuracy of the MACE-Omat models at about two orders of magnitude higher measured throughput. The most compact reduces the energy, force and stress errors of the fastest existing universal MLIP by 61.4%, 48.1% and 34.3% at 1.92 times its saturated throughput. All five variants complete multimillion-atom simulations on a single GPU and run molecular dynamics for 2.048 billion atoms on 1,024 16-GB NVIDIA V100 GPUs at 83.3--91.2% weak-scaling efficiency. Compared with the MEAM empirical potential, DPA4C-Nano reaches 1.8 and 2.5 times the saturated throughput in single-GPU scans on the same V100 hardware for diamond carbon and FCC copper, respectively. DPA4C therefore brings quantum-trained universal accuracy into a regime of speed and system size previously associated with empirical potentials.