arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LEAP:面向GPU内核生成的代码强化学习的自适应剪枝轻量环境反馈框架

LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation

Tankun Li, Zhi Chen, Yaohua Tang

arXiv 2608.01804首次发表:更新:

发表机构

Moore Threads(摩尔线程)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LEAP是针对GPU内核生成的代码强化学习框架,通过难度条件剪枝与基于排名的奖励公式解决低级代码RL的信号稀疏性等问题,收敛更快且性能更优。

AI 中文摘要

通过强化学习(RL)对大型语言模型(LLM)进行训练后,代码生成能力已显著提升。为规避评判网络的高内存占用,当前最先进的框架采用无评判范式,如结合基于规则的验证沙箱的分组相对策略优化(GRPO)。然而,将这些框架应用于CUDA内核生成等低级系统编程时面临严峻挑战:二元通过/失败奖励导致严重的信号稀疏性,多轮环境反馈循环则存在编译延迟过高及轨迹间奖励稀释问题。本研究提出LEAP(Lean Environment-Feedback via Adaptive Pruning),一种针对低级硬件加速器适配优化的可扩展、计算高效的多轮RL框架。LEAP包含难度条件剪枝(DCP),一种动态门控机制,可自适应从多轮扩展中剔除简单及过度灾难性任务,仅将资源密集型编译与硬件探索聚焦于高价值复杂任务。为无需手动超参数工程即可充分实现这些路径,我们提出基于排名的奖励公式,通过从GRPO回滚组内的成对竞赛结果推导无尺度相对优势,该方法固有地惩罚简单提示下的令牌低效性,同时最大化挑战性分布上的学习梯度。实证评估显示,LEAP实现了更优的首轮熟练度与稳健的多轮调试恢复力,且收敛速度快于未剪枝的多轮基线,为低级代码RL建立了实用范式。

英文摘要

Post-training large language models (LLMs) via reinforcement learning (RL) has significantly advanced code generation capabilities. To bypass the heavy memory footprint of critic networks, current state-of-the-art frameworks leverage critic-free paradigms like Group Relative Policy Optimization (GRPO) tied to rule-based verification sandboxes. However, applying these frameworks to low-level systems programming, such as CUDA kernel generation-presents severe challenges: binary pass/fail rewards introduce severe signal sparsity, while multi-turn environmental feedback loops suffer from prohibitive compilation latencies and reward dilution across trajectories. In this work, we introduce LEAP (Lean Environment-Feedback via Adaptive Pruning), a scalable and computationally efficient multi-turn RL framework optimized for low-level hardware accelerator alignment. LEAP features Difficulty-Conditioned Pruning (DCP), a dynamic gating mechanism that adaptively cuts off simple and overly catastrophic tasks from multi-turn expansion, focusing resource-heavy compilation and hardware exploration exclusively on high-value, complex tasks. To fully operationalize these paths without manual hyperparameter engineering, we propose a Rank-Based Reward formulation. By deriving scale-free relative advantages from pairwise tournament outcomes within the GRPO rollout group, our method inherently penalizes token inefficiency on simple prompts while maximizing learning gradients on challenging distributions. Empirical evaluations show that LEAP achieves superior first-turn proficiency and robust multi-turn debugging resilience while converging faster than unpruned multi-turn baselines, establishing a practical paradigm for low-level code RL.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑