arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33074cs.LGcs.SE

KernelZero:协同进化生成器与编码器以持续改进GPU内核生成

KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation

Changxin Ke, Rui Zhang, Zixiang Fang, Zhenghong Li, Yuanbo Wen, Jiashuo Shen, Shuo Wang, Jiaming Guo, Ling Li, Qi Guo, Yunji Chen

AI总结:

提出KernelZero协同进化框架,通过生成器与编码器交替优化及CA-GRPO策略,在保证正确性后优化性能,显著提升GPU内核生成效率,超越多个基线模型。

AI中文摘要:

高性能GPU内核对于现代机器学习系统至关重要,然而自动生成既正确又高效的内核仍然具有挑战性。现有的基于LLM的方法面临两个主要限制:与模型当前能力对齐的高质量训练数据稀缺,以及内核正确性与性能之间的固有权衡。为应对这些挑战,我们提出了KernelZero,一个协同进化框架,通过两个专门模型持续改进GPU内核生成:一个生成器(Proposer),从API集合生成Torch模块,以及一个编码器(Coder),将其转换为CUDA或Triton内核。KernelZero使用前沿驱动的模块生成机制,基于编码器当前的弱点持续产生能力对齐的训练模块。它进一步引入了正确性感知的组相对策略优化(CA-GRPO),仅在正确性足够可靠后才优化性能。通过交替优化生成器和编码器,KernelZero形成了一个自动课程,实现有针对性的、训练高效的性能提升。实验上,KernelZero-7B在CUDA上超越了Claude-4.5-Sonnet,在Triton上超越了DeepSeek-V4-Pro。在KernelBench Level 1和2上,它分别实现了CUDA pass@1得分75.8%和69.6%,pass@10达到100%和97%。在Triton上,它分别实现了pass@1得分77.2%和72.5%。

英文摘要:

High-performance GPU kernels are essential to modern machine learning systems, yet automatically generating kernels that are both correct and efficient remains challenging. Existing LLM-based approaches face two major limitations: the scarcity of high-quality training data aligned with the model's current capabilities, and the inherent trade-off between kernel correctness and performance. To address these challenges, we propose KernelZero, a co-evolution framework that continuously improves GPU kernel generation through two specialized models: a Proposer that generates Torch modules from API sets and a Coder that translates them into CUDA or Triton kernels. KernelZero uses a frontier-driven module generation mechanism to continuously produce capability-aligned training modules based on the Coder's current weaknesses. It further introduces Correctness-Aware Group Relative Policy Optimization (CA-GRPO), which optimizes performance only after correctness becomes sufficiently reliable. By alternating the optimization of the Proposer and Coder, KernelZero forms an automatic curriculum that enables targeted and training-efficient capability improvement. Empirically, KernelZero-7B surpasses Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton. On KernelBench Level 1 and 2, it achieves CUDA pass@1 scores of 75.8% and 69.6%, respectively, with pass@10 reaching 100% and 97%. On Triton, it achieves pass@1 scores of 77.2% and 72.5%, respectively.

补充信息

↑