迈向科学:用于科学计算的人工智能芯片探索
Ascend to Science: Exploration of AI Chips for Scientific Computing
浏览论文内容
中文总结 AI 辅助
研究探讨人工智能芯片用于科学计算的条件,以昇腾910 NPU为平台刻画瓶颈,针对五个应用研究开发特定映射,通过协调数值公式、执行位置和数据移动,使原生人工智能NPU实现数值稳健性、性能和可扩展性,区分优化原则与实现细节。
中文摘要 AI 辅助
面向人工智能的加速器迅速崛起,重塑了以低精度张量引擎为核心的计算系统,这给高性能计算社区带来了一个实际问题:在何种条件下,这种硬件能够支持需要数值稳健性、不规则内存访问和可扩展性的科学工作负载?我们以昇腾910 NPU系列为代表性张量中心平台,刻画了阻碍科学代码直接部署的精度、执行和内存层次瓶颈。然后,我们针对五个应用研究——HPL-MxP、LRSVD、SGEMM-cube、PQSim和SMC-X,开发并评估了特定于工作负载的映射,结合了异构执行、混合精度数值公式、精度仿真、分层内存编排和通信-计算重叠。这些研究表明,当数值公式、执行位置和数据移动得到协调处理时,原生人工智能NPU可以实现数值稳健性、有竞争力的性能和令人满意的可扩展性。我们的结果提供了一个实践案例研究,展示了科学工作负载如何适应以张量为中心的架构,同时区分可转移的优化原则和昇腾特定的实现细节。
英文摘要
The rapid rise of AI-oriented accelerators has reshaped compute systems around low-precision tensor engines, raising a practical question for the HPC community: under what conditions can such hardware support scientific workloads that demand numerical robustness, irregular memory access, and scalability? Using the Ascend 910 NPU series as a representative tensor-centric platform, we characterize precision, execution, and memory-hierarchy bottlenecks that hinder the direct deployment of scientific codes. We then develop and evaluate workload-specific mappings across five application studies -- HPL-MxP, LRSVD, SGEMM-cube, PQSim, and SMC-X -- combining heterogeneous execution, mixed-precision numerical formulations, precision emulation, hierarchical memory orchestration, and communication--computation overlap. These studies show that AI-native NPUs can achieve numerical robustness, competitive performance, and satisfactory scalability when numerical formulation, execution placement, and data movement are addressed in a coordinated manner. Our results provide a state-of-the-practice case study of how scientific workloads can be adapted to tensor-centric architectures, while distinguishing transferable optimization principles from Ascend-specific implementation details.