发表机构
Princeton University; Institute for Research and Innovation in Software for High Energy Physics (IRIS-HEP)(普林斯顿大学; 高能物理软件研究与创新研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对 HL-LHC 时代 AMD GPU 需求,提出 Rust 内核引擎 rawkward,通过优化模式恢复 CUDA 级性能,实现供应商无关的高能物理分析内核。
AI 中文摘要
高亮度大型强子对撞机(HL-LHC)将要求分析吞吐量实现数量级的提升,而这些提升日益需要来自非单一供应商的 GPU。El Capitan、Frontier 和 LUMI 等领先级系统基于 AMD 加速器构建,然而 Scikit-HEP 分析栈——尤其是 Awkward Array——一直以 CUDA 优先发展。我们报告了 $rawkward$,一个由 Rust 支持的内核引擎,为 Awkward Array 的嵌套、不规则、可变长度数据结构添加了 ROCm/HIP 后端。我们的核心发现是,将 CUDA 内核直接源码级移植到 HIP 会导致不规则内核性能损失 5 到 10 倍,因为 AMD 的 64 通道波前、更高的寄存器压力和更昂贵的分歧行为与 NVIDIA 的 32 线程束有根本性不同。我们展示了一小组可复用的优化模式——循环展平、128 位向量化加载、拆分融合内核以及基于配置文件的启动配置——在不改变公共 API 的情况下恢复了 CUDA 级性能。一个 Rust 宏和匹配调度层保持单一、后端无关的调用点,同时发出特定于供应商的内核策略,并且类型系统在编译时强制执行缓冲区大小和生命周期正确性。在一个双插槽 AMD Instinct MI210 节点上,我们测量到 GPU 相对于 128 个 EPYC 7763 CPU 内核的加速比从 1.03 倍(带宽受限的 $sum$)到 12.5 倍($count$)不等,并且 Rust CPU 内核在总体上匹配或击败了现有的 C++ 内核(十二个内核的几何平均运行时间比为 0.37 倍)。我们认为这些模式构成了一种实用的、可移植性能的、供应商无关的 HEP 分析内核的配方。
英文摘要
The High-Luminosity LHC (HL-LHC) will demand order-of-magnitude gains in analysis throughput, and increasingly those gains must come from GPUs that are not made by a single vendor. Leadership-class systems such as El Capitan, Frontier and LUMI are built on AMD accelerators, yet the Scikit-HEP analysis stack---and Awkward Array in particular---has grown up CUDA-first. We report on $rawkward$, a Rust-backed kernel engine that adds a ROCm/HIP backend for Awkward Array's nested, jagged, variable-length data structures. Our central finding is that a naive source-level port of CUDA kernels to HIP loses $5$--$10\times$ in performance on irregular kernels, because AMD's $64$-lane wavefronts, higher register pressure and more expensive divergence behave fundamentally differently from NVIDIA's $32$-thread warps. We show that a small, reusable set of optimization patterns---loop flattening, $128$-bit vectorized loads, splitting fused kernels, and profile-guided launch configuration---recovers CUDA-class performance without changing the public API. A Rust macro-and-match dispatch layer keeps a single, backend-agnostic call site while emitting vendor-specific kernel strategies, and the type system enforces buffer-size and lifetime correctness at compile time. On a two-socket AMD Instinct MI210 node we measure GPU speedups from $1.03\times$ (bandwidth-bound $sum$) up to $12.5\times$ ($count$) over $128$ EPYC~7763 CPU cores, and the Rust CPU kernels match or beat on aggregate the incumbent C++ kernels (geometric-mean runtime ratio $0.37\times$ across twelve kernels). We argue that these patterns constitute a practical recipe for performance-portable, vendor-agnostic HEP analysis kernels.
Comments8 pages, 2 figures