发表机构
Shanghai Key Laboratory of Scalable Computing and Systems; School of Computer Science, Shanghai Jiao Tong University(上海可扩展计算与系统重点实验室; 上海交通大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HIERA是一种跨实现空间的GPU内核优化分层规划框架,在KernelBench实验中优于无训练方法,还在科学计算模板算子上实现1.53倍加速,无需额外训练即可接近CUDA-L1性能。
AI 中文摘要
高性能GPU内核是现代深度学习与科学计算的基础。随着工作负载日益多样化,且GPU硬件快速迭代,开发高效的自动化GPU内核生成与优化方法愈发重要。现有基于大语言模型(LLM)的方法通常在固定实现空间内优化,这要么限制了优化灵活性,要么降低了搜索效率。我们提出HIERA,这是一种用于GPU内核优化的分层搜索空间规划框架。HIERA构建带有契约增强的任务规格说明,从PyTorch算子、CUDA库和自定义CUDA内核中选择合适的实现空间,并利用性能剖析反馈与专家知识指导结构化迭代优化。在KernelBench上针对多种工作负载水平和基础LLM开展的实验显示,HIERA相比现有无训练方法,在整体实现有效性、样本效率和优化性能上表现更优,且无需额外模型训练即可与基于训练的CUDA-L1相媲美。针对科学计算中一种专用模板算子的案例研究进一步显示,其相比cuDNN实现了1.53倍的加速,证明了该通用框架在标准机器学习工作负载之外的潜力。
英文摘要
High-performance GPU kernels underpin modern deep learning and scientific computing. As workloads become increasingly diverse and GPU hardware evolves rapidly, developing efficient methods for automated GPU kernel generation and optimization has become increasingly important. Existing LLM-based methods typically optimize within a fixed implementation space, limiting either optimization flexibility or search efficiency. We propose \textsc{HIERA}, a hierarchical search-space planning framework for GPU kernel optimization. \textsc{HIERA} constructs contract-augmented task specifications, selects an appropriate implementation space across PyTorch operators, CUDA libraries, and custom CUDA kernels, and uses profiling feedback and expert knowledge to guide structured iterative refinement. Experiments on KernelBench across multiple various workload levels and base LLMs show that \textsc{HIERA} delivers stronger overall implementation validity, sample efficiency, and optimization performance than existing training-free methods, while remaining competitive with the training-based CUDA-L1 without additional model training. A case study on a specialized stencil operator from scientific computing further achieves a \(1.53\times\) speedup over cuDNN, demonstrating the potentiality of the general framework beyond standard machine-learning workloads.