发表机构
Spacial Intelligence Labs(空间智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究对比了CUDA C++、Rust、Triton在哈希阻塞GPU工作负载上的性能,发现规则阶段三者性能接近,非规则阶段Triton慢一个数量级,Rust接近CUDA,还分析了性能差异原因并修复了cuda-oxide的一个缺陷。
AI 中文摘要
GPU语言对比几乎总是在平铺密集线性代数上进行,此时所有工具链表现相近,差异很小。我们用CUDA C++、通过NVIDIA的cuda-oxide实现的Rust以及Triton,实现了相同的哈希阻塞TSDF融合内核,并在具有相反特性的工作负载上进行了测量:该工作负载采用开放寻址哈希表,带有比较交换插入、依赖于线程的数据探测深度以及竞争散射。结果呈现分化:在遍历截断带并进行累加的规则阶段,三种语言的性能差距很小;而在进行探测与插入的非规则阶段,Rust与手写CUDA C++性能接近,Triton则慢了一个数量级以上。在通常被基准测试的工作上,语言选择几乎无影响;而在未被关注的工作上,语言选择的代价很高。我们将两种性能差距归因于各语言无法表达的特定特性,而非性能比值。Triton的代价源于必须运行到编译时绑定的探测循环,以及tl.atomic_cas不支持掩码,这迫使它使用了CUDA中没有的暂存结构。Rust的代价在所有指令计数中都不可见:其内核在相同占用率下发出的指令更少、比较交换操作更少、寄存器更少,但速度更慢。硬件计数器显示问题出在L1驻留上:GPU范围的原子加载必须在流式多处理器(SM)间保持一致性,而NVIDIA的L1缓存不具备该特性,因此读取共享位置的类型正确方式是每次访问都绕过缓存。Triton的绑定探测还对融合存在正确性问题:在普通深度轨迹达到的负载因子下,它会静默丢弃块,导致重建丢失部分表面。我们还报告了在cuda-oxide中发现并修复的一个缺陷,现已合并到上游:其作用域内的原子加载和存储在生成实际内核的构建模式下根本无法调用。
英文摘要
GPU language comparisons are almost always run on tiled dense linear algebra, where every toolchain is good and the differences are small. We implement the same hash-blocked TSDF fusion kernel in CUDA C++, in Rust through NVIDIA's cuda-oxide, and in Triton, and measure it on a workload with the opposite character: an open-addressed hash table with compare-exchange insertion, data-dependent per-lane probe depth, and contended scatter. The result is a split. On the regular stage, which walks a truncation band and accumulates, all three languages land within a small factor of each other. On the irregular stage, which probes and inserts, Rust stays close to hand-written CUDA C++ while Triton is more than an order of magnitude slower. Language choice is nearly free on the work that is usually benchmarked and expensive on the work that is not. We attribute both gaps to specific things the languages cannot express, not to ratios. Triton's cost follows from a probe loop that must run to a compile-time bound and from tl.atomic_cas taking no mask, which forces a scratch structure with no counterpart in CUDA. Rust's cost was invisible in every instruction count: its kernel issues fewer instructions, fewer compare-exchanges and fewer registers at identical occupancy, yet was slower. Hardware counters located it in L1 residency. A GPU-scope atomic load must be coherent across SMs, no NVIDIA L1 is, so the type-correct way to read a shared location bypasses the cache on every access. Triton's bounded probe is also a correctness problem for fusion: at load factors an ordinary depth trajectory reaches, it silently discards blocks and the reconstruction loses patches of surface with nothing reported. We also report a defect found and fixed in cuda-oxide itself, now merged upstream: its scoped atomic load and store could not be called at all in the build mode that produces real kernels.
Comments23 pages, 5 figures, 4 tables. Includes a correctness result for TSDF fusion implementations: at hash load factors reached by ordinary depth trajectories, the Triton implementation silently discards blocks. Code, raw measurement CSVs and an interactive viewer: https://github.com/realitymatrix/what-irregularity-costs