AI 中文总结
MakoXC是一种模块化矩阵对齐XC计算引擎,通过三项关键技术重构稀疏性,使XC计算提速显著,可在64个GPU上5分钟内完成泛素的端到端DFT计算。
AI 中文摘要
密度泛函理论(Density Functional Theory,DFT)是材料科学与药物发现领域不可或缺的工具,但交换关联(exchange-correlation,XC)计算因立方级复杂度成为主要性能瓶颈。尽管线性标度方法利用电子近视性降低渐近复杂度,但其生成的不规则稀疏工作负载会隐藏隐式稀疏性,无法高效利用现代AI加速器。本文提出MakoXC,这是一种模块化的矩阵对齐XC计算引擎,将近视性诱导的稀疏性重构为适合加速器的规则计算。MakoXC协同设计了三项关键技术:(1)矩阵对齐单元(Matrix-Aligned Cells)将近视性诱导的相互作用重组为密集、适配加速器的数据簇;(2)稀疏性引导激活(Sparsity-Guided Activation)将更深层的隐式稀疏性转换为数值正确的结构化执行,实现实用的线性标度;(3)内核融合流水线(Kernel-Fused Pipeline)将碎片化工作负载整合为统一的计算密集型执行路径,充分释放加速器吞吐量。大量评估显示,MakoXC相较于标准XC计算平均提速67.8倍,相较于最先进的线性标度方法平均提速4.7倍。当集成到生产级商用DFT软件包中时,MakoXC可在64个GPU上将XC计算扩展到泛素(1231个原子,def2-SVP基组),使端到端DFT计算在5分钟内完成。通过将XC计算重构为统一的结构化计算,MakoXC证明了科学工作负载如何在最大化AI加速器并行效率的同时实现真正的低复杂度。
英文摘要
Density Functional Theory (DFT) is indispensable for materials science and drug discovery, yet the exchange--correlation (XC) evaluation remains a major bottleneck due to its cubic scaling. Although linear-scaling methods exploit electronic nearsightedness to reduce asymptotic complexity, they produce irregular sparse workloads that hide implicit sparsity and prevent efficient use of modern AI accelerators. We present MakoXC, a modular matrix-aligned XC evaluation engine that rearchitects nearsightedness-induced sparsity into regular, accelerator-friendly computations. MakoXC co-designs three key techniques: (1) Matrix-Aligned Cells reorganize nearsightedness-induced interactions into dense, accelerator-aligned data clusters; (2) Sparsity-Guided Activation translates deeper implicit sparsity into numerically correct structured execution for practical linear scaling; and (3) Kernel-Fused Pipeline consolidates fragmented workloads into a unified, compute-intensive execution path that fully unleashes accelerator throughput. Extensive evaluations show that MakoXC achieves average speedups of 67.8$\times$ speedup over standard XC evaluation and 4.7$\times$ over state-of-the-art linear-scaling methods. When integrated into a production-grade commercial DFT package, MakoXC scales XC evaluation to ubiquitin (1,231 atoms, def2-SVP) on 64 GPUs, enabling the end-to-end DFT calculation to complete in under five minutes. By restructuring XC evaluation into a unified, structured computation, MakoXC demonstrates how scientific workloads can achieve genuine low complexity while maximizing parallel efficiency on AI accelerators.
CommentsAccepted in the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC'26)