arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22964cs.DC

在Tenstorrent RISC-V加速器上探索谱元方法

Exploring spectral element methods on the Tenstorrent RISC-V accelerator

Daniyal Arshad, Nick Brown

首次发表
浏览论文内容

中文总结 AI 辅助

本文将HPC迷你应用Nekbone的AX内核移植到Tenstorrent Wormhole RISC-V加速器,解决数据转置瓶颈后性能优于24核至强Platinum CPU且功耗显著降低。

中文摘要 AI 辅助

商用RISC-V硬件的日益普及激发了其在高性能计算(HPC)领域的应用兴趣,PCIe加速器卡为其近期落地提供了可行途径。Tenstorrent Wormhole是这类产品之一,它在128个Tensix核心上配备了专用的向量和矩阵单元,且应用广泛。本文探索将Nekbone的AX内核移植到Wormhole加速器上,Nekbone是源自戈登·贝尔奖获奖谱元求解器Nek5000的一款广泛使用的HPC迷你应用,该内核用于求解泊松算子,我们描述了该算法在Tensix核心上的映射方式。初始性能结果显示,z方向梯度计算所需的主机侧数据转置是严重瓶颈。为此,我们研究了两种优化策略,取得了显著改进:在128个Tensix核心上处理100000个单元时,达到242.97 GFLOPS的性能,优于24核至强Platinum CPU,且功耗约降低7倍。

英文摘要

The growing availability of commodity RISC-V hardware has sparked interest in its use for High Performance Computing (HPC), with PCIe accelerator cards offering a practical near-term pathway to adoption. The Tenstorrent Wormhole is one example, with dedicated vector and matrix units across 128 Tensix cores, and is widely available. In this paper, we explore porting the AX kernel of Nekbone, a widely used HPC mini-application derived from the Gordon Bell Prize-winning Nek5000 spectral element solver, onto the Wormhole accelerator. This kernel evaluates the Poisson operator, and we describe the mapping of the algorithm onto the Tensix. The initial performance results reveal that the host-side data transposition, required for the z-direction gradient computation, is a severe bottleneck. Consequently, we investigated two optimisation strategies that yield dramatic improvements, achieving 242.97 GFLOPS for 100000 elements across 128 Tensix cores, outperforming a 24-core Xeon Platinum CPU and drawing approximately 7 times less power.

补充信息

↑