arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21058cs.DCcs.AI

LLM生成的GPU内核实际上能达到真实工作负载的多少?

How Much of a Real Workload Can LLM-Generated GPU Kernels Actually Reach?

Gaurav Agarwal, Ashish Garg, Isha Singhal

AI总结:

本研究评估LLM生成GPU内核在真实工作负载中的实际收益,发现可寻址时间比例有限,并引入DLRM-Bench基准,同时揭示KernelBench正确性检查的漏洞。

AI中文摘要:

语言模型现在可以编写性能优于PyTorch的GPU内核。我们在KernelBench级别1上评估了五种模型配置,发现一个前沿模型在91.1%的问题上生成了正确的内核,并在56个问题中的22个上独立验证了加速效果,包括三个卷积,中位加速比为1.235倍。开放权重模型远远落后:最好的模型达到30.4%的正确率,只有三个验证的加速,并且解决了零个卷积。然后我们提出了文献中未涉及的问题:这样的内核控制了多少真实模型墙钟时间的比例?对三个领域的七个工作负载进行剖析,我们发现可寻址比例范围从8.9%到58.2%。在Transformer上,80-86%的运行时间花费在cuBLAS GEMM和FlashAttention上,将现实端到端改进限制在约1%,并且该比例随模型规模增大而缩小。在推荐系统上,该比例为58.2%,集中在单个嵌入内核中。我们引入了DLRM-Bench,即KernelBench格式的12个推荐系统内核问题,并测量了41.7%的胜率,中位加速比为1.552倍,预计端到端改进为8.63%。另外,我们展示了KernelBench的正确性检查(使用绝对容差的此HTTP URL)在60个级别1问题中的4个上被全零张量满足。我们自己的结果中有两个内核在检测到之前利用了这一点,其中一个得分283倍的内核只写入了其输出缓冲区的0.3%。我们提出了尺度不变的替代方案,并发布了所有879个评估结果。

英文摘要:

Language models can now write GPU kernels that outperform PyTorch. We evaluate five model configurations on KernelBench level 1 and find that a frontier model produces correct kernels for 91.1% of problems and independently verified speedups on 22 of 56, including three convolutions, with a median of 1.235x. Open-weights models are far behind: the best reaches 30.4% correct with three verified speedups and solves zero convolutions. We then ask a question the literature does not: what fraction of a real model's wall clock do such kernels govern? Profiling seven workloads across three domains, we find the addressable fraction ranges from 8.9% to 58.2%. On transformers, 80-86% of runtime is spent in cuBLAS GEMM and FlashAttention, bounding realistic end-to-end improvement at roughly 1%, and the fraction shrinks with model scale. On recommenders it is 58.2%, concentrated in a single embedding kernel. We introduce DLRM-Bench, 12 recommender kernel problems in KernelBench format, and measure a 41.7% win rate at a 1.552x median there, projecting 8.63% end-to-end. Separately, we show that KernelBench's correctness check (torch.allclose with an absolute tolerance) is satisfied by a tensor of zeros on 4 of 60 level-1 problems. Two kernels in our own results exploited this before we detected them, including one scored at 283x that wrote 0.3% of its output buffer. We propose scale-invariant replacements and release all 879 evaluations.

补充信息

相关深度报道

↑