AI 中文总结
研究将GPU内核的静态PTX指标与运行时配对,发现结构错误在静态信号中明确,语义错误则不可见。表明静态PTX增量门可分离结构与语义变化,测量的运行时增量受硬件和噪声影响,不能替代。
AI 中文摘要
我们将每个GPU内核的静态PTX指标(寄存器、溢出、指令计数)与五种GPU类型(RTX 3060、A10、L40S、A100 SXM4和H100 NVL)上的CUDA事件定时运行时进行配对。在这个语料库和工具链中,静态信号和测量信号在一个轴上清晰分离。对于任何给定的(正确、有缺陷)对,每对之间的Delta-regs和Delta-instrs在所有五个GPU上都是相同的。测量的Delta-perf%则不然。改变内核工作的结构错误在静态信号中是明确的。gelu_triton_buggy变体删除了一个前导0.5因子,减少了8条指令和8个寄存器。在RTX 3060上相应的测量Delta-perf%为+3.2%,在这些语料库内核所占据的亚毫秒尺度的运行间噪声范围内。用另一个常量替换一个常量的语义错误对静态信号是不可见的。softmax_triton_buggy变体在掩码加载时用other=0.0替换-inf,编译后的PTX字节相同。本文的有限主张是,对于这个语料库和工具链,静态PTX增量门是一个可移植的预过滤器,可将结构变化与语义变化分开;在这个尺度上测量的运行时增量对硬件和噪声敏感,不能替代。
英文摘要
We pair each GPU kernel's static PTX metrics (registers, spills, instruction count) with CUDA-event-timed runtime on five GPU classes: RTX 3060, A10, L40S, A100 SXM4, and H100 NVL. In this corpus and toolchain the static and measured signals separate cleanly along one axis. Per-pair Delta-regs and Delta-instrs are identical across all five GPUs for any given (correct, buggy) pair. Measured Delta-perf% is not. Structural bugs that change the kernel's work are unambiguous in the static signal. The gelu_triton_buggy variant, which drops a leading 0.5 factor, removes 8 instructions and 8 registers. The corresponding measured Delta-perf% on RTX 3060 is +3.2%, within the run-to-run noise band at the sub-millisecond scale these corpus kernels occupy. Semantic bugs that swap one constant for another are invisible to the static signal. The softmax_triton_buggy variant, which substitutes other=0.0 for -inf on the masked load, compiles to byte-identical PTX. The paper's bounded claim is that, for this corpus and toolchain, a static-PTX delta gate is a portable pre-filter that separates structural from semantic changes; measured runtime deltas at this scale are hardware- and noise-sensitive and are not a substitute.
Comments9 pages, 2 figures, LNCS format