arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

正确但缓慢:现代领域特定语言中GPU内核评估差距的实证研究

Correct but Slow: An Empirical Study of the GPU Kernel Evaluation Gap in Modern Domain-Specific Languages

Tingxi Li, Ravishka Rathnasuriya, Wei Yang

arXiv 2607.04454首次发表:更新:

AI 中文总结

研究现代GPU领域特定语言内核评估的正确性与性能差距。通过对特定内核研究,发现基于正确性评估会接纳严重减速内核,原因因内核家族而异,提出两项轻量级检查作为补充筛选标准。

AI 中文摘要

现代GPU领域特定语言(如Triton和TileLang)越来越多地用于实现专用深度学习内核及作为自动内核生成系统的目标语言。现有DSL内核评估通过基于参考的数值验证来确定正确性,但对替换质量未作考量。我们使用来自五个算子类别的22个Triton和TileLang内核进行研究,得出三项结果。

英文摘要

Modern GPU domain-specific languages (DSLs), such as Triton and TileLang, are increasingly used to implement specialized deep-learning kernels and as target languages for automated kernel-generation systems. Existing DSL-kernel evaluations establish correctness through reference-based numerical validation -- necessary, but silent on replacement quality: a functionally valid kernel may still fall far below the throughput of the optimized library operator it is intended to replace. We study this correctness-performance gap using 22 Triton and TileLang kernels from five operator categories on NVIDIA A100 and GH200 GPUs, asking whether correctness-based evaluation identifies kernels unsuitable as library replacements, why such failures occur, and how they can be detected without exhaustive benchmark coverage. The study yields three results. \emph{First}, correctness-based evaluation can admit severe slowdowns: an idiomatic TileLang LayerNorm kernel passes KernelBench's correctness check while running more than 300$\times$ slower than the PyTorch baseline. \emph{Second}, the causes differ by kernel family. TileLang normalization and reduction slowdowns are mainly repairable authoring defects, such as sequential reductions and unnecessary dtype conversions, whereas convolution and large general matrix multiplication (GEMM) retain residual gaps after optimization due to code-generation and autotuning-coverage limits; vendor-library algorithm selection contributes only marginally. \emph{Third}, two lightweight checks -- library-relative efficiency and roofline utilization -- are complementary screening criteria: together they flag every functionally valid but inefficient kernel in our suite and separate repairable authoring defects from structural residuals.

Comments17 pages, 2 figures --- Jul. 6th EDIT: update artifact link in the paper --- Jul. 13th EDIT: correct errors in reference, add new reference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑