发表机构
Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自然语言生成CUDA内核的专业壁垒与奖励黑客问题,本文提出CUDA-Harness框架,通过中间结构化生成、基于合成的验证及反馈自适应进化实现内核生成优化,且泛化性良好。
AI 中文摘要
开发高性能CUDA内核需要算法实现、正确性验证及硬件感知并行优化的专业知识,形成了显著的专业壁垒,使得从自然语言生成CUDA内核(Text2CUDA)成为刚需。同时,大型语言模型(LLM)的通用代码生成能力推动了一系列基于LLM的CUDA内核生成研究,这些研究主要聚焦于从PyTorch等高级框架到CUDA的转译(Torch2CUDA),而非Text2CUDA,后者要求模型理解高级输入语义并处理低级内核实现与验证。此外,这些方法因依赖预定义测试输入而易受奖励黑客攻击。本文提出CUDA-Harness框架,用于利用智能体从自然语言生成与优化CUDA内核。具体而言,引入中间结构化生成(Intermediate-Structured Generation)以连接高级语义理解与低级内核生成;为缓解Text2CUDA中的奖励黑客攻击,构建基于合成的验证(Synthesis-Based Verification)以提供隔离测试数据与渐进式验证;进一步提出反馈自适应进化(Feedback-Adaptive Evolution),一种优先保证正确性同时优化性能的内核进化策略。最后,通过大量实验证明了CUDA-Harness的有效性,进一步评估显示其在不同LLM、硬件平台及C转译CUDA(C-to-CUDA)任务上具备泛化能力。
英文摘要
Developing high-performance CUDA kernels demands specialized knowledge in algorithm implementation, correctness validation, and hardware-aware parallel optimization, creating a substantial expertise barrier and making generating CUDA kernels directly from natural language (Text2CUDA) essential. Meanwhile, the general-purpose code generation capability of Large Language Models (LLMs) prompts a series of works exploring LLM-based CUDA kernel generation. They mainly focus on transpilation from high-level frameworks such as PyTorch to CUDA (Torch2CUDA) rather than Text2CUDA, where models must understand the high-level input semantics and handle low-level kernel implementation and validation. Additionally, these methods are vulnerable to reward hacking due to reliance on predefined test inputs. In this paper, we propose CUDA-Harness, a framework for harnessing agentic CUDA kernel generation and optimization from natural language. Specifically, we introduce Intermediate-Structured Generation to connect high-level semantic understanding with low-level kernel generation. To dilute reward hacking in Text2CUDA, we construct Synthesis-Based Verification to provide isolated test data and progressive validation. Furthermore, we propose Feedback-Adaptive Evolution, a kernel evolution strategy that prioritizes correctness while optimizing performance. Finally, through extensive experiments, we demonstrate the effectiveness of CUDA-Harness, with further evaluations illustrating generalization across LLMs, hardware platforms, and to C-to-CUDA transpilation.