发表机构
Huawei Research(华为研究)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出面向异构超级云的计算模型与ReDSEa工具链,自动化映射、负载均衡等,在鲲鹏920与昇腾910系统上实现CH分解和SVD分析高达17倍和88倍的加速。
AI 中文摘要
异构超级云系统通过在不同资源间实现可扩展且高效的任务分配,正在变革数据分析领域。本文提出了专门用于在异构超级云上加速稠密Cholesky(CH)分解和基于奇异值分解(SVD)的分析的计算模型与性能估计技术。我们的ReDSEa工具链自动化了映射、负载均衡、调度、并行性和计算重叠。在配备华为鲲鹏920 ARM CPU和昇腾910 AI加速器的异构系统上实现后,我们的LLVM编译器工具链采用了针对递归、迭代和分块计算的新型性能模型,相较于经过全面优化的48核CPU实现,CH实现了高达17倍的加速,SVD实现了88倍的加速。
英文摘要
Heterogeneous Supercloud systems are transforming data analytics by enabling scalable and efficient task distribution across diverse resources. This paper presents computational models and performance estimation techniques tailored for accelerating Dense Cholesky (CH) Decomposition and Singular Value Decomposition (SVD)-based analytics on Heterogeneous Superclouds. Our ReDSEa tool-chain automates mapping, load balancing, scheduling, parallelism, and computation overlap. Implemented on a heterogeneous system with a Huawei Kunpeng 920 ARM CPU and an Ascend 910 AI accelerator, our LLVM compiler tool-chain employs novel performance models for recursive, iterative, and blocked computations, achieving up to 17x speedup for CH and 88x for SVD over fully optimized 48-core CPU implementations.