arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于代码优化的强化学习

Reinforcement Learning for Code Optimization

Pierre Chambon, Kunhao Zheng, Juliette Decugis, Benoit Sagot, Gabriel Synnaeve

arXiv 2607.25970首次发表:更新:

发表机构

FAIR at Meta; Inria(Meta FAIR研究院; 法国国家信息与自动化研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对代码优化中强化学习的问题,通过构建DMC - Optim和校准沙盒、组合正确性与速度及调整GRPO和评估等方法,使执行时间可学习,在DMC - Optim等任务上取得显著效果,提升了代码优化能力。

AI 中文摘要

强化学习用于代码正确性已确立:让模型生成程序,针对隐藏测试用例运行并奖励通过的解决方案。将其扩展到代码优化看似简单,只需在奖励中加入执行时间。但实践中,一旦时间驱动奖励,测量噪声、奖励稀疏性或GRPO不稳定性等小问题就会使强化学习失败。我们通过三个阶段使执行时间可学习:一是通过构建DMC - Optim和校准沙盒来测试代码;二是在强化学习环境中组合正确性和速度并使用离线模拟器预测最有前景的配置来将速度转化为奖励;三是通过调整GRPO和评估以适应更稀疏、有噪声的定时执行设置让模型从奖励中学习。在DMC - Optim上,最强的优化感知配置在Qwen 2.5 7B上使严格的前50% pass@1从18.0%提高到31.3%,在CWM 32B上从30.7%提高到50.4%。在更严格的百分位数如前30%时收益进一步增加,同时保持纯正确性分数。当定时沙盒退化时,稳健优化强化学习比标准RLVR提高100%到200%。在LCB上,CWM 32B在与标准RLVR的中位数样本速度比较中赢得高达83%。相对于每个问题最快的正确人类提交,它达到人类复杂度类改进率的约一半。

英文摘要

RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and reward solutions that pass. Extending this to code optimization seems straightforward: just add execution time to the reward. But in practice, once timing drives the reward, small problems in measurement noise, reward sparsity, or GRPO instability overwhelm the signal and make RL fail: generated solutions are barely faster, and more of them can fail. We make execution time learnable through three stages: (1) how code is tested, by building DMC-Optim with large optimization tests and a calibrated sandbox; (2) how speed is turned into reward, by composing correctness and speed in the RL environment and using an offline simulator to predict the most promising configurations; and (3) how the model learns from that reward, by adapting GRPO and evaluation to the sparser, noisier timed-execution setting. On DMC-Optim, the strongest optimization-aware configurations improve strict top-50% pass@1 from 18.0% to 31.3% on Qwen 2.5 7B and from 30.7% to 50.4% on CWM 32B. These gains further increase at stricter percentiles such as top-30%, with 125% relative improvement for CWM 32B, while preserving pure-correctness scores. When the timing sandbox is degraded, robust optimization RL reaches 100% to 200% improvement over standard RLVR, depending on the evaluation criterion. On LCB, CWM 32B wins up to 83% of median-sample speed comparisons against standard RLVR. Relative to the fastest correct human submissions per problem, it reaches about half the human rate of complexity-class improvements (13% vs. 22%).

Comments126 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑