发表机构
University at Albany (SUNY); Université Paris-Est Créteil(纽约州立大学奥尔巴尼分校; 巴黎东克雷泰伊大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文在真实HPC集群上验证MoA推导的Transformer内核内存最优成本函数,发现并修复GPU原子争用回归(达2.5倍加速),揭示拓扑导致的成本差异,并提出固定DNF的ONF重写方法论。
AI 中文摘要
我们验证了通过数组数学(MoA)推导出的Transformer内核的内存最优成本函数。配套论文I-IV形式化推导了注意力前向、反向、融合前向+反向、解码以及完整模块(RMSNorm、门控MLP)的内核,作为硬件无关的规范(DNF),通过gamma转换为机器特定的实现(ONF),并与PyTorch进行了机器精度级别的验证。本文在CPU和GPU上,针对两个HPC集群(Purdue Anvil、NCSA Delta)的实测性能检验了这些预测。有三项结果尤为突出。(1)我们发现并修复了一个GPU回归:融合前向+反向,虽被证明可避免物化O(n^2)中间结果,但最初在GPU上因原子争用而运行速度慢于朴素实现。性能分析确认原子指令数量多出2.00倍;针对性的ONF重写扭转了这一局面,实现了高达2.5倍的加速。(2)相同的推导因拓扑不同而产生显著不同的实际成本:在一个集群上出现535倍的NUMA局部性惩罚,而在另一个集群上仅为不到3倍的超额订阅,这表明最优部署是机器数组结构的函数。(3)我们报告了一个部分解决的异常:相同的指称计算在CPU上C语言比Fortran运行更快,但在GPU上Fortran比C语言更快,该问题被缩小到一个主导内核和一种内存延迟停顿机制(3.17倍的时间差距与3.35倍的停顿差距相匹配)。我们将硬件特定优化视为常规的ONF重写,并保持固定的、已验证的DNF,这是一种将AI扩展到不断演进的硬件上而无需重新推导正确性的候选方法论。
英文摘要
We validate memory-optimal cost functions for transformer kernels derived via the Mathematics of Arrays (MoA). Companion Papers I-IV formally derive kernels for attention forward, backward, fused forward+backward, decode, and the complete block (RMSNorm, gated MLP) as a hardware-independent specification (DNF) transformed to a machine-specific realization (ONF) via gamma, with verification to machine precision against PyTorch. This paper checks those predictions against measured performance on two HPC clusters (Purdue Anvil, NCSA Delta) across CPU and GPU. Three results stand out. (1) We identify and fix a GPU regression: fusing forward+backward, proven to avoid materializing an O(n^2) intermediate, initially ran slower than naive on GPU due to atomic contention. Profiling confirmed 2.00x more atomic instructions; a targeted ONF rewrite reversed it, yielding up to 2.5x speedup. (2) Identical derivations produce markedly different real costs by topology: 535x NUMA-locality penalty on one cluster vs <3x oversubscription on another, showing optimal deployment is a function of the machine's array structure. (3) We report a partially resolved anomaly: identical denotational computations run faster in C than Fortran on CPU but faster in Fortran than C on GPU, narrowed to one dominant kernel and one memory-latency stall mechanism (3.17x time gap matches 3.35x stall gap). We treat hardware-specific optimization as a routine ONF rewrite with fixed, verified DNF, a candidate methodology for scaling AI onto evolving hardware without re-deriving correctness.
CommentsCompanion validation for Papers I-IV. Paper I arXiv:2606.07713, Paper II HAL:05659212, Paper III arXiv:2607.19456, Paper IV HAL:05734881 (https://hal.science/hal-05734881). 26 pages, 20 figures (placeholders in v1, real figures in v2). Implementation: https://github.com/womenflyplanes/moa-attention-verified-mullin. Allocation CIS261396 on Purdue Anvil and NCSA Delta via ACCESS