arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

驯服 GPU 内核中的位级行为与张量核心:黑盒重构、编译器强制与静态验证

Taming Bitwise Behavior in GPU Kernels with Tensor Core: Black-Box Reconstruction, Compiler Enforcement, and Static Verification

Ziteng Yang, Nicholas J. Riasanovsky, Warren Deng, Vivek Sarkar

arXiv 2609.11356首次发表:更新:

发表机构

Georgia Institute of Technology; Meta(佐治亚理工学院; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对 GPU 内核位级行为不确定性问题,提出归约顺序描述符、黑盒重构、编译器强制平衡树归约及静态位等价检查器,实现与 cuBLAS 位级匹配并提升性能。

AI 中文摘要

机器学习系统对 GPU 内核的确定性和数值可复现性要求日益提高,然而同一内核的确定性实现仍可能逐位不同。浮点归约顺序是主要原因,此外还有部分和精度、融合乘加运算和舍入位置等因素。这些选择可能由手工编码、由诸如 Triton 的块级语言选定,或隐藏在诸如 cuBLAS 或 rocBLAS 的闭源库中。为性能而选择的 tile 形状因此也决定了算术运算,可能破坏批次不变性。保持固定顺序可能造成高达 20% 的性能开销,而自动调优器无法识别哪些配置是位等价的。我们刻画了决定归约和通用矩阵乘法(GEMM)位级行为的因素。首先,我们引入了一种 GEMM 归约顺序的描述符,包括 split-K GEMM 中 K 维的划分。利用该描述符,我们首次对闭源库的算术进行了黑盒重构,以实现位级正确性。我们的 Triton GEMM 系列在 Blackwell 和 Hopper 上的所有测试案例中均与 NVIDIA cuBLAS 匹配。对于具有融合后置操作的现实 LLM 形状,其性能匹配或超过 cuBLAS。其次,我们在 Triton 降级过程中强制平衡树归约,并引入一种数据布局优化,使 GB300 和 H100 上 27 个内核中的 19 个性能达到自由顺序性能的 10% 以内。第三,我们开发了用于编译后 GPU 内核之间位等价性的可靠静态检查器,包括首个覆盖 NVIDIA PTX 和 AMD GCN 的检查器。集成到 Triton 的自动调优器中后,该检查器将搜索限制在单个位等价类内。

英文摘要

Determinism and numerical reproducibility are increasingly required of GPU kernels in machine learning systems, yet deterministic implementations of the same kernel can still differ bit for bit. Floating-point reduction order is the primary cause, alongside partial-sum precision, fused multiply-add operations, and rounding placement. These choices may be hand-coded, selected by a block-level language such as Triton, or hidden inside a closed-source library such as cuBLAS or rocBLAS. A tile shape chosen for performance therefore also determines the arithmetic, potentially breaking batch invariance. Preserving a fixed order can cost up to 20 percent, while an autotuner cannot identify which configurations are bitwise equivalent. We characterize the factors determining the bitwise behavior of reductions and general matrix multiplication (GEMM). First, we introduce a descriptor of GEMM reduction order, including the partitioning of K in split-K GEMM. Using it, we perform the first black-box reconstruction of a closed-source library's arithmetic for bit-level correctness. Our family of Triton GEMMs matches NVIDIA cuBLAS in all tested cases on Blackwell and Hopper. For realistic LLM shapes with fused epilogues, it matches or exceeds torch.compile performance. Second, we enforce balanced-tree reduction during Triton lowering and introduce a data-layout optimization that brings 19 of 27 kernels on GB300 and H100 within 10 percent of free-order performance. Third, we develop sound static checkers for bitwise equivalence between compiled GPU kernels, including the first checker spanning NVIDIA PTX and AMD GCN. Integrated into Triton's autotuner, the checker restricts search to a single bit-equivalence class.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑