arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用可变块稀疏矩阵库加速原子模拟

Accelerating Atom Simulations with Variable-Block Sparse Matrix Library

Zhanghao Zhouyin, Hong Guo

arXiv 2609.04397首次发表:更新:

AI 中文总结

该研究针对原子模拟中稀疏块矩阵的局部结构未被利用的问题,提出保留可变原子块的分布式稀疏矩阵库VBCSR,其通过分组同形块并分派至优化内核实现加速,在基准测试及InP纳米粒子应用中表现优于参考实现。

AI 中文摘要

现代原子模拟越来越多地采用局域轨道来表示量子算子,产生的稀疏块矩阵其块形状随化学物种和基组选择而变化。传统标量稀疏格式单独存储每个块的元素,掩盖了这种局部结构并限制了高效块算法的使用。我们提出VBCSR,一种分布式稀疏矩阵库,可保留可变大小的原子块并加速大规模原子模拟的核心线性代数运算。统一接口自动将标量、均匀基组和多物种算子映射到压缩稀疏行(CSR)、块压缩稀疏行(BSR)或可变块压缩稀疏行(VBCSR)。我们的高级加速方法将相同形状的块分组并将其分派到优化的密集内核。在报告的基准测试中,VBCSR在几个块稀疏基准测试中优于测试的可通过Python访问的参考实现。我们进一步在包含超过10^6个原子的InP纳米粒子应用中演示了VBCSR。

英文摘要

Modern atomistic simulations increasingly employ localized orbitals to represent quantum operators, yielding sparse block matrices whose block shapes vary with chemical species and basis choice. Conventional scalar sparse formats store the entries of each block individually, obscuring this local structure and limiting the use of efficient block algorithms. We present VBCSR, a distributed sparse matrix library that preserves variable-size atomic blocks and accelerates the core linear algebra of large-scale atomistic simulations. A unified interface automatically maps scalar, uniform-basis, and multispecies operators to compressed sparse row (CSR), block sparse row (BSR), or variable-block compressed sparse row (VBCSR). Our advanced acceleration method groups blocks of equal shape and dispatches them to optimized dense kernels. In the reported benchmarks, VBCSR outperforms the tested Python-accessible reference implementations for several block-sparse benchmarks. We further demonstrate VBCSR in an InP nanoparticle application containing more than \(10^6\) atoms.

Comments8 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑