arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05709physics.chem-ph

cuSOAP:原子位置平滑重叠描述符的GPU加速生成器

cuSOAP: a GPU-accelerated Generator of Smooth Overlap of Atomic Positions Descriptor

Hongyu Yan, Yuqing Xia, Yong Wei, Minghan Chen, Hanning Chen

首次发表
浏览论文内容

中文总结 AI 辅助

cuSOAP是GPU加速的SOAP描述符生成器,基于PyTorch和融合内核,比DScribe快两个数量级,支持百万原子体系,实现近线性扩展和高效并行。

中文摘要 AI 辅助

原子位置平滑重叠(SOAP)描述符是分子机器学习中最广泛采用的原子环境表示之一,但评估其及其导数的成本仍然是基于SOAP的原子间势的主要瓶颈,特别是在复杂、多物种凝聚相环境所需的大径向和角向基组尺寸下。我们提出了cuSOAP,一个用于原子级SOAP向量及其解析导数的GPU加速生成器。cuSOAP基于PyTorch构建,通过融合的CUDA和Triton内核,评估高斯型轨道和多项式径向基的投影系数及其笛卡尔梯度的闭式表达式或求积,从而消除了朴素张量公式中的多吉字节中间体。该软件包是CPU参考实现DScribe的直接替代品,重现了其构造函数签名、特征排序和输出,精度达到约10^{-6},并直接接受原子模拟环境(ASE)的Atoms对象作为结构输入。在NVIDIA Grace-Blackwell(GB200)节点上,单个Blackwell GPU为1000分子水团簇生成完整的描述符加雅可比矩阵工作负载,比DScribe快达两个数量级,即在整个(n_max, l_max)超参数映射上加速13倍至227倍,且加速比随角向带宽限制l_max的增加而增长。仅描述符生成在多达百万原子水团簇上扩展为t ∝ n^1.34,接近线性,在一个GPU上处理耗时19.5秒,通过共享内存多进程驱动程序在节点的四个GPU上处理耗时5.7秒,展示了约90%的并行效率,且随着设备增加无饱和迹象。这些结果使得大规模凝聚相模拟中的实时SOAP评估成为可能。

英文摘要

The Smooth Overlap of Atomic Positions (SOAP) descriptor is one of the most widely adopted representations of atomic environments in molecular machine learning, but the cost of evaluating it and its derivatives remains a principal bottleneck of SOAP-based interatomic potentials, particularly at the large radial and angular basis sizes demanded by complex, multi-species condensed-phase environments. We present cuSOAP, a GPU-accelerated generator of atom-wise SOAP vectors and their analytic derivatives. Built on PyTorch, cuSOAP evaluates closed-form expressions or quadratures for the projection coefficients and their Cartesian gradients for Gaussian-type-orbital and polynomial radial bases, through fused CUDA and Triton kernels that eliminate the multi-gigabyte intermediates of a naive tensor formulation. The package is a drop-in replacement for the CPU-based reference DScribe, reproducing its constructor signature, feature ordering, and output to within ${\sim}10^{-6}$, and accepts structures directly as Atomistic Simulation Environment (ASE) Atoms objects. On an NVIDIA Grace--Blackwell (GB200) node, a single Blackwell GPU generates the full descriptor-plus-Jacobian workload for a 1000-molecule water cluster up to two orders of magnitude faster than DScribe, i.e., from $13\times$ to $227\times$ across the entire $(n_{\max}, l_{\max})$ hyperparameter map, with the speedup growing with the angular band limit $l_{\max}$. Descriptor-only generation scales as $t \propto n^{1.34}$, close to linear, up to a million-atom water cluster, which is processed in 19.5s on one GPU and, through a shared-memory multiprocessing driver, in 5.7s on the four GPUs of the node, demonstrating $\sim$90% parallel efficiency with no sign of saturation as devices are added. These results bring on-the-fly SOAP evaluation for large-scale condensed-phase simulation within reach.

发表机构

  • Heriot-Watt University(赫瑞-瓦特大学)
  • Chinese University of Hong Kong(香港中文大学)
  • University of North Georgia(北乔治亚大学)
  • Wake Forest University(维克森林大学)
  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

↑