用小内存占用的机器学习原子间势解锁多组分块体材料的分子动力学模拟
Unlocking Multi-Component Bulk-Materials Molecular Dynamics with a Small-Footprint Machine Learning Interatomic Potential
浏览论文内容
中文总结 AI 辅助
本文提出一种HBM占用仅为现有MLIPs不到3%的机器学习原子间势,用144个NVIDIA A100 GPU实现了含1.14×10^9个原子的6组分块体系统的分子动力学模拟,突破了多组分块体MD模拟的算力限制。
中文摘要 AI 辅助
与纳米材料不同,块体材料需要在大空间尺度(约10^9个原子或更多)上进行分子动力学(MD)模拟,才能充分捕捉其原子尺度的物理性质。此前,机器学习原子间势(MLIPs)的引入将MD模拟扩展到了该尺度,但即使是单组分块体系统,在高端超级计算机上也需要数万个GPU。然而,多组分块体MD模拟几乎无法实现,因为现有MLIPs的高带宽内存(HBM)占用——即使对于单组分系统已经很大——在多组分场景下会爆炸式增长。本文提出了一种HBM占用小的MLIP,其占用量不到现有MLIPs的3%,仅用数百个GPU即可解锁多组分块体MD模拟。这是通过首先确定特征向量和中间张量是现有MLIPs中HBM占用的两个主要贡献者来实现的;为解决这两个来源,通过引入物理和化学知识降低了特征向量的维度,并通过将所有内核激进地融合为一个巨型内核消除了中间张量。在评估中,所提出的MLIP使用144个NVIDIA A100 GPU对包含1.14×10^9个原子的6组分块体系统进行了MD模拟,而此前这种MD模拟的空间尺度仅限于单组分系统,且通常需要配备数万个GPU的高端超级计算机才能实现。
英文摘要
Bulk materials, as opposed to nanomaterials, require molecular dynamics (MD) simulations on a large spatial scale (~10^9 atoms or more) to adequately capture their atomic-scale physical properties. Previously, the introduction of machine-learning interatomic potentials (MLIPs) has extended MD to this scale, but even single-component bulk systems require tens of thousands of GPUs on high-end supercomputers. However, multi-component bulk MD simulations remain barely achievable, as the HBM footprint of existing MLIPs - already substantial for single-component systems - grows explosively in multi-component scenarios. This paper proposes an MLIP with a small HBM footprint - less than 3% that of existing MLIPs - unlocking multi-component bulk MD using only hundreds of GPUs. This is achieved by first identifying feature vectors and intermediate tensors as the two primary contributors to HBM footprints in existing MLIPs. To address these two sources, the dimensionality of the feature vectors has been reduced by introducing physical and chemical knowledge, and intermediate tensors have been eliminated by aggressively fusing all kernels into a single mega-kernel. In evaluation, the proposed MLIP has used 144 NVIDIA A100 GPUs to perform MD simulations on a 6-component bulk system with 1.14x10^9 atoms, while previously such MD simulation spatial scale has been restricted to unary systems and typically achieved on high-end supercomputers equipped with tens of thousands of GPUs.