arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17379cs.CLcs.AI

PTXBench:用于结合架构特定PTX优化GPU内核的大型语言模型基准测试与适配

PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization with Architecture-specific PTX

  • Carnegie Mellon University(卡内基梅隆大学)
  • RadixArk
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Genghan Zhang, Yixin Dong, Chengze Fan, Zhichen Zeng, Yueming Yuan, Shaowei Zhu, Kunle Olukotun

中文总结 AI 辅助

研究提出PTXBench基准测试,评估LLM结合架构特定PTX优化GPU内核的能力,发现现有模型能力不均衡,经微调Qwen3.6-27B可改善部分任务但泛化仍不均,该基准为相关研究提供可审计测试平台。

中文摘要 AI 辅助

我们提出了PTXBench,这是一个用于评估和适配大型语言模型(LLM)以使用架构特定PTX进行GPU内核优化的基准测试。PTXBench在H100和B200 GPU上的GEMM和注意力工作负载中,测量功能正确性、所选目标指令是否在运行时执行,以及相较于前沿库的加速比。我们的评估显示,架构特定PTX的能力仍不均衡:在复杂的注意力反向工作负载上成功率大幅下降,执行目标指令并不一定能转化为有竞争力的性能;所有评估模型均未在整个测试套件中始终匹配前沿库。我们进一步使用监督微调适配Qwen3.6-27B,修复条件训练改善了多项任务,但泛化能力仍不均衡;除数据集规模外,数据覆盖范围、平衡性及推理教师的质量也至关重要。PTXBench提供了一个可审计的测试平台,用于测量和提升LLM利用不断演进的GPU架构的能力。

英文摘要

We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven: success rates fall substantially on complex attention backward workloads, and executing the target instructions does not necessarily translate into competitive performance. No evaluated model consistently matches frontier libraries across the suite. We further adapt Qwen3.6-27B using supervised fine-tuning. Repair-conditioned training improves several tasks, but generalization remains uneven; data coverage, balance, and the quality of the reasoning teacher matter in addition to dataset size. PTXBench provides an auditable testbed for measuring and improving LLMs' ability to exploit evolving GPU architectures.

↑