KernelZero:协同进化生成器与编码器以持续改进GPU内核生成
KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation
AI总结:
提出KernelZero协同进化框架,通过生成器与编码器交替优化及CA-GRPO策略,在保证正确性后优化性能,显著提升GPU内核生成效率,超越多个基线模型。
AI中文摘要:
高性能GPU内核对于现代机器学习系统至关重要,然而自动生成既正确又高效的内核仍然具有挑战性。现有的基于LLM的方法面临两个主要限制:与模型当前能力对齐的高质量训练数据稀缺,以及内核正确性与性能之间的固有权衡。为应对这些挑战,我们提出了KernelZero,一个协同进化框架,通过两个专门模型持续改进GPU内核生成:一个生成器(Proposer),从API集合生成Torch模块,以及一个编码器(Coder),将其转换为CUDA或Triton内核。KernelZero使用前沿驱动的模块生成机制,基于编码器当前的弱点持续产生能力对齐的训练模块。它进一步引入了正确性感知的组相对策略优化(CA-GRPO),仅在正确性足够可靠后才优化性能。通过交替优化生成器和编码器,KernelZero形成了一个自动课程,实现有针对性的、训练高效的性能提升。实验上,KernelZero-7B在CUDA上超越了Claude-4.5-Sonnet,在Triton上超越了DeepSeek-V4-Pro。在KernelBench Level 1和2上,它分别实现了CUDA pass@1得分75.8%和69.6%,pass@10达到100%和97%。在Triton上,它分别实现了pass@1得分77.2%和72.5%。
英文摘要:
High-performance GPU kernels are essential to modern machine learning systems, yet automatically generating kernels that are both correct and efficient remains challenging. Existing LLM-based approaches face two major limitations: the scarcity of high-quality training data aligned with the model's current capabilities, and the inherent trade-off between kernel correctness and performance. To address these challenges, we propose KernelZero, a co-evolution framework that continuously improves GPU kernel generation through two specialized models: a Proposer that generates Torch modules from API sets and a Coder that translates them into CUDA or Triton kernels. KernelZero uses a frontier-driven module generation mechanism to continuously produce capability-aligned training modules based on the Coder's current weaknesses. It further introduces Correctness-Aware Group Relative Policy Optimization (CA-GRPO), which optimizes performance only after correctness becomes sufficiently reliable. By alternating the optimization of the Proposer and Coder, KernelZero forms an automatic curriculum that enables targeted and training-efficient capability improvement. Empirically, KernelZero-7B surpasses Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton. On KernelBench Level 1 and 2, it achieves CUDA pass@1 scores of 75.8% and 69.6%, respectively, with pass@10 reaching 100% and 97%. On Triton, it achieves pass@1 scores of 77.2% and 72.5%, respectively.