发表机构
Iowa State University; Cisco AI Research(爱荷华州立大学; 思科人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对 CUDA 内核生成,提出 CudaPerf 框架,结合可验证执行奖励与结构代码感知奖励,分离线排序和在线训练两阶段,利用执行反馈迭代细化,通过实验证明其显著优于基线,在加速和正确性上有大幅提升。
AI 中文摘要
具有可验证奖励的强化学习(RLVR)已成为增强大语言模型推理能力以进行优化代码生成的强大技术。然而,现有 RLVR 方法主要依赖基于结果的信号,忽略了对生成优化代码至关重要的程序性能关键结构属性。本文提出 CudaPerf,这是一个反射式 RL 框架,结合了可验证执行奖励和从并行化特征派生的结构代码感知奖励。CudaPerf 分两个阶段运行:离线成对排序模块和在线 RL 训练阶段。为进一步增强学习,利用执行反馈进行迭代细化。还引入了一个数据集。通过多个基准评估发现,CudaPerf 显著优于强基线,在加速方面分别实现高达 5 倍和 3.32 倍的提升,在正确性方面分别提高了 17%和 7%。
英文摘要
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful technique to enhance the reasoning capacity of LLMs for optimized code generation. However, existing RLVR approaches primarily rely on outcome-based signals such as correctness and speedup, overlooking performance-critical structural properties of programs that are essential for generating optimized code. In this work, we propose CudaPerf, a reflective RL framework that incorporates both verifiable execution rewards and structural code-aware rewards derived from parallelization features (e.g., memory coalescing, occupancy, Arithmatic Intensity, and synchronization patterns). CudaPerf operates in two stages: (1) an offline pairwise ranking module that learns to distinguish strong and weak program candidates via contrastive comparisons, and (2) an online RL training phase that jointly optimizes for correctness, performance, and structural efficiency through a unified reward signal. To further enhance learning, CudaPerf utilizes iterative refinement using execution feedback enabling progressive improvement of generated candidates. We also introduce a dataset comprising 2.9k C to CUDA and 1k PyTorch to CUDA programs, each paired with diverse input configurations and multiple CUDA implementations encompassing diverse optimization strategies. CudaPerf is evaluated across multiple benchmarks comprising both C to CUDA and PyTorch to CUDA transformations. Empirical findings suggest that CudaPerf significantly outperforms strong baselines, including Qwen-3-32B (for C to CUDA) and CUDA Agent (for PyTorch to CUDA) by achieving up to 5X & 3.32X improvements in speedup, and 17% & 7% improvements in correctness, respectively.